4 Identification & search
Read Introduction first.
4.1 Database search (DDA)
In data-dependent acquisition (DDA), peptide identification is typically performed by comparing experimental tandem mass spectra with theoretical spectra generated from a reference protein database (McLafferty 1981; Eng et al. 1994; Tyanova et al. 2016). Reference protein sequence databases are commonly obtained from UniProt (The UniProt Consortium 2018). The protein sequences are first digested in silico according to the same enzymatic rules used during sample preparation, producing a list of theoretical peptides. For each peptide, a theoretical fragmentation spectrum is generated by predicting the fragment ions expected during tandem mass spectrometry.
The experimental MS2 or MS/MS spectra are then compared against these theoretical spectra in a process known as database searching. Specialized search engines, including Andromeda (Cox et al. 2011) implemented in MaxQuant (Cox and Mann 2008), SEQUEST (Eng et al. 1994), MSFragger (Kong et al. 2017), and MS-GF+ (Kim and Pevzner 2014), evaluate how well each theoretical spectrum matches the observed experimental spectrum. Each comparison produces a peptide-spectrum match (PSM), which is assigned a score reflecting the quality of the match (Sinitcyn et al. 2018). Although the underlying scoring algorithms differ between search engines, they all aim to distinguish correct peptide identifications from incorrect matches.
4.1.1 Proteomics databases
TODO: Add content on proteomics databases (e.g., UniProt, sequence database formats and construction).
4.1.2 Rescoring strategies and scoring algorithms
TODO: Add content on rescoring strategies and scoring algorithms (e.g., Percolator, machine learning-based rescoring).
4.2 Database independent acquisition (DIA search)
TODO: Add content on the DIA search approach (e.g., spectral libraries, DIA-specific search engines).
4.3 de novo sequencing
TODO: Add content on de novo sequencing (e.g., algorithms, scoring, limitations).
4.4 Handling multiple search engines and combining results
TODO: Add content on combining results from multiple search engines.
4.4.1 Comparison of search engines and search modes
TODO: Add content comparing search engines and search modes.
4.5 FDR calculation and statistical validation
Peptide-spectrum matches (PSMs) are assigned scores that reflect how well an experimental spectrum matches a theoretical spectrum generated from a protein sequence database. Although these scores rank candidate peptide identifications, they do not directly indicate the probability that a match is correct. Consequently, an additional statistical framework is required to estimate the number of true and false peptide identifications.
The most widely used approach for estimating identification confidence is the target-decoy strategy, which provides an estimate of the false discovery rate (FDR) (Elias and Gygi 2007; Benjamini and Hochberg 1995; Storey and Tibshirani 2003). In this approach, a decoy database is generated alongside the target protein database, typically by reversing the protein sequences. The resulting decoy sequences preserve key properties of the target database, such as amino acid composition and sequence length, but do not correspond to real proteins. Consequently, any PSM assigned to a decoy sequence is considered a false positive.
The numbers of target and decoy identifications are then used to estimate the proportion of false positive identifications among the target matches. This estimation relies on the assumption that incorrect matches are equally likely to occur against the target and decoy databases (Elias and Gygi 2007). By selecting a score threshold corresponding to a desired FDR, commonly 1%, researchers can control the expected proportion of false positive peptide identifications while retaining as many true identifications as possible.
4.5.1 Entrapment databases for FDR estimation
TODO: Add content on entrapment databases for FDR estimation.
4.6 Protein inference
TODO: Add content on protein inference (e.g., parsimony principle, protein groups, shared peptides).
4.7 PTM analysis
4.7.1 PTM identification
Identification of post-translational modifications (PTMs) in mass spectrometry-based proteomics relies primarily on detecting mass shifts that correspond to the addition or removal of a specific chemical group on a modifiable amino acid residue (Lennon and Walsh 1999). D uring database searching, these expected mass shifts are incorporated into the search space, allowing modified peptide candidates to be matched to experimental spectra.
In addition to characteristic mass shifts, many PTMs produce diagnostic fragment ions or neutral losses during peptide fragmentation, providing additional evidence for modification assignment. For example, phosphorylated peptides may generate the phosphotyrosine-specific immonium ion at m/z = 216.042, which serves as a diagnostic reporter ion for phosphotyrosine-containing peptides (Olsen et al. 2007). Phosphorylated peptides can also undergo neutral loss of phosphoric acid during fragmentation, producing characteristic fragment ions that further support phosphorylation site identification (Olsen et al. 2007). These diagnostic ions and neutral loss patterns can be incorporated into the search space of database search engines to improve peptide identification (Cox et al. 2011). For example, the Andromeda search engine explicitly evaluates spectra both with and without the expected neutral loss at the modification site and retains the higher of the two scores as the final peptide-spectrum match score (Cox et al. 2011).
4.7.2 PTM localisation
TODO: Add content on PTM localisation (e.g., localisation scores such as Ascore/PTM score, site ambiguity).
4.8 Specific search strategies
There are multiple data types and acquisitions methods that require specialized search strategies. These include cross-linking mass spectrometry, glycoproteomics, top-down proteomics, proteogenomics and metaproteomics.
4.8.1 Cross-linking mass spectrometry
TODO: Add content on cross-linking mass spectrometry (e.g., search engines, scoring, FDR estimation).
4.8.2 Glycoproteomics
TODO: Add content on glycoproteomics (e.g., search engines, scoring, FDR estimation).
4.8.3 Top-down proteomics
TODO: Add content on top-down proteomics (e.g., search engines, scoring, FDR estimation).
4.8.4 Proteogenomics
TODO: Add content on proteogenomics (e.g., search engines, scoring, FDR estimation).
4.8.5 Metaproteomics
TODO: Add content on metaproteomics.