A Face in the Crowd: Recognizing Peptides Through Database Search
Peptide identification via tandem mass spectrometry sequence database searching is a key method in the array of tools available to the proteomics researcher. The ability to rapidly and sensitively acquire tandem mass spectrometry data and perform peptide and protein identifications has become a commonly used proteomics analysis technique because of advances in both instrumentation and software. Although many different tandem mass spectrometry database search tools are currently available from both academic and commercial sources, these algorithms share similar core elements while maintaining distinctive features. This review revisits the mechanism of sequence database searching and discusses how various parameter settings impact the underlying search. Peptide identification via tandem mass spectrometry sequence database searching is a key method in the array of tools available to the proteomics researcher. The ability to rapidly and sensitively acquire tandem mass spectrometry data and perform peptide and protein identifications has become a commonly used proteomics analysis technique because of advances in both instrumentation and software. Although many different tandem mass spectrometry database search tools are currently available from both academic and commercial sources, these algorithms share similar core elements while maintaining distinctive features. This review revisits the mechanism of sequence database searching and discusses how various parameter settings impact the underlying search. Innovations in tandem mass spectrometry (MS/MS) 1The abbreviations used are: MS/MStandem MSPTMpost-translational modification. have enabled the rapid growth of proteomics for chemists and biologists alike. In addition to the evolution of the instrument hardware to acquire spectra more rapidly and sensitively, improvements to data analysis facilitates the process of identifying and quantifying peptides and proteins for a wide user base. Researchers new to the field can benefit from a review of the fundamentals of tandem mass spectrometry sequence database searching, and even experienced proteomics researchers may profit from considerations of practical strategies to maximize information extraction from data. tandem MS post-translational modification. In a typical shotgun proteomics experiment, the sample is first denatured to enable proteolysis. An enzyme such as trypsin is applied to cleave proteins to smaller peptide components. This peptide mixture is usually then separated on a liquid chromatography (LC) column; more complex samples may necessitate prior fractionation by strong cation exchange or isoelectric focusing. Peptides are subjected to ionizing voltage as they are electrosprayed from the column, and these ions are introduced into the near-vacuum environment of the mass spectrometer. In a tandem mass spectrometry experiment, the instrument will first acquire a survey or precursor scan, also referred to as an MS scan, which measures all intact peptide ions eluting into the mass spectrometer at that given time. One or more peptide ions are selected, sequentially isolated, fragmented, and the resulting fragment ions are measured to produce an MS/MS spectrum. This process is repeated to automatically acquire MS/MS spectra on as many different peptide ions as possible throughout the LC gradient. See Fig. 1 for a schematic showing the relationship between precursor ions in the MS scans and the resulting MS/MS scans. While the MS/MS spectrum contains the peptide fragmentation pattern, the experimental peptide mass and charge state are obtained from the precursor ion measure in the MS spectrum. To clarify terminology, the intact peptide ions in the MS scans are termed precursor ions. Fragment ions are also called product ions and MS/MS spectra are referred to as product ion spectra. MS and MS/MS spectra can also referred to as MS1 and MS2 spectra, respectively. Although collision-induced dissociation (CID) is the most common way to generate product ions, others including electron capture dissociation (ECD), electron transfer dissociation (ETD), and infrared multiphoton dissociation (IRMPD) have also been developed. After data acquisition, the MS/MS spectra are searched against a protein sequence database to identify the underlying peptides and proteins represented in the acquired spectra. A 1994 paper by Mann and Wilm introduced the strategy of searching protein sequence databases using peptide sequence tags interpreted from MS/MS spectra (1Mann M. Wilm M. Error-tolerant identification of peptides in sequence databases by peptide sequence tags.Anal. Chem. 1994; 66: 4390-4399Crossref PubMed Scopus (1318) Google Scholar). These tags of 3 to 5 consecutive amino acids were inferred by recognizing chains of amino acid mass differences between peaks. The masses flanking the tag along with the inferred sequence were then matched to peptide sequences drawn from a database of protein sequences. This approach was shown to be capable of identifying peptides with post-translational modifications (PTMs) by allowing for differences between the measured peptide mass and the calculated mass of peptides in the database. At first, tag inference was done manually so it was a relatively slow process that required some domain expertise to perform. In recent years, automated systems for inferring sequence tags have been joined by new tools to reconcile mass differences between spectra and sequences to yield new capabilities in modification and mutation identification (2Tanner S. Shu H. Frank A. Wang L.C. Zandi E. Mumby M. Pevzner P.A. Bafna V. InsPecT: identification of posttranslationally modified peptides from tandem mass spectra.Anal. Chem. 2005; 77: 4626-4639Crossref PubMed Scopus (504) Google Scholar, 3Liu C. Yan B. Song Y. Xu Y. Cai L. Peptide sequence tag-based blind identification of post-translational modifications with point process model.Bioinformatics. 2006; 22: e307-313Crossref PubMed Scopus (29) Google Scholar, 4Tabb D.L. Ma Z.Q. Martin D.B. Ham A.J. Chambers M.C. DirecTag: accurate sequence tags from peptide MS/MS through statistical scoring.J. Proteome Res. 2008; 7: 3838-3846Crossref PubMed Scopus (100) Google Scholar). Uninterpreted MS/MS database search algorithms, also introduced in 1994, rapidly became the standard method for protein identification. Eng et al. described the SEQUEST algorithm to match MS/MS spectra to peptides drawn from protein sequence databases (5Eng J.K. McCormack A.L. Yates 3rd, J.R. An approach to correlate tandem mass spectral data of peptides with amino acid sequences in a protein database.J. Am. Soc. Mass Spectr. 1994; 5: 976-989Crossref PubMed Scopus (5472) Google Scholar). This approach gained the ability to identify post-translational modifications the following year (6Yates 3rd, J.R. Eng J.K. McCormack A.L. Schieltz D. Method to correlate tandem mass spectra of modified peptides to amino acid sequences in the protein database.Anal. Chem. 1995; 67: 1426-1436Crossref PubMed Scopus (1109) Google Scholar). With no requirement of manual interpretation of spectra prior to running searches, the ability to automate the search process increases throughput and opened the door for nonexperts to perform the analyses. Subsequent years saw the publication of both commercial and open-source tools that implemented this approach. For a detailed review of specific search tools and a collection of search engine comparisons, please see (7Kapp E. Schütz F. Overview of tandem mass spectrometry (MS/MS) database search algorithms.in: Current Protocols in Protein Science. John Wiley & Sons, Inc., 2007Google Scholar, 8Nesvizhskii A.I. Protein identification by tandem mass spectrometry and sequence database searching.Methods Mol. Biol. 2007; 367: 87-119PubMed Google Scholar, 9Brosch M. Swamy S. Hubbard T. Choudhary J. Comparison of Mascot and X!Tandem performance for low and high accuracy mass spectrometry and the development of an adjusted Mascot threshold.Mol. Cell. Proteomics. 2008; 7: 962-970Abstract Full Full PubMed Scopus Google Scholar, A. H. of MS/MS search algorithms for analysis of spectra from electron transfer dissociation Chem. PubMed Scopus Google Scholar, The of mass data and search algorithm on peptide identification in Chem. 2007; PubMed Scopus Google Scholar, Schütz F. D. Eng J.K. An and accurate of available MS/MS search and 2005; 5: PubMed Scopus Google Scholar, The of ions on search algorithm performance for dissociation PubMed Scopus Google Scholar, M. M. D. V. E. J. of the using ion mass spectrometry and search Proteome Res. PubMed Scopus Google Scholar). The common elements of these tools and practical considerations are the specific of this a most MS/MS database search tools perform the all a collection of MS/MS spectra, a sequence database to peptides of the these peptides against the experimental spectra, and peptide A schematic of process is in Fig. For MS/MS an experimental peptide mass can be from the precursor and or measured precursor charge A database search will peptides from the sequence database that are of the mass as the experimental peptide mass for the spectrum. The of peptides that against spectrum are referred to as The of the precursor mass is by the accuracy of the mass used to measure the precursor spectra. The of peptides is also by such as the enzyme and post-translational modifications in the search. A of fragment ion masses are then calculated for peptide The method used to fragment the peptides which of fragment ion masses are and for collision-induced dissociation and and for electron transfer These fragment ions are against in the experimental spectrum and a or is peptide is and the peptides for spectrum are usually in are the of protein and peptide sequences for database search all possible sequences of these tools search to the sequences in the database. researchers will a sequence database that contains all proteins for the of from This database can be by sequences of common researchers also these databases by of sequences to enable search strategy for in protein identifications by mass 2007; PubMed Scopus Google Scholar). tandem mass spectra are for peptides database search algorithms the of in peptides from protein sequences. In the of this and are by a The enzyme parameter the of peptides that are and against experimental spectrum. The impact of this is search possible peptide to MS/MS spectrum. the peptide sequences that have masses the precursor ion be with a given MS/MS spectrum. This precursor mass may be in the or on the mass used to generate the data and available in a search the enzyme parameter this precursor mass parameter to the of peptides that with experimental spectrum. a peptide and an experimental MS/MS the database search algorithm how a sequence to the spectrum. the peptide sequence from a of is a for spectra that and are peaks. The resulting search are used for the all the peptides with a MS/MS spectrum such that the peptide at the of the these will be used the search to the of peptides from the of spectra subjected to the search. a these strategies have from a of algorithms for calculated fragment ion masses against experimental MS/MS spectra have from the of and The used in SEQUEST measures the to which experimental and spectra the of between the spectrum and a spectrum from the The of this can be by a calculated by spectrum to the in The is the with no the calculated from a of This for the spectrum is by a between the peptides termed the it is calculated by the between the by the Mascot protein identification by searching sequence databases using mass spectrometry PubMed Scopus Google introduced a statistical of match the of fragment ions in the the the of the spectrum a and the of peptide sequences with the spectrum. have been used to the of a identification to a on the of fragment that between and spectra. These the used by L. Xu M. mass spectrometry search Proteome Res. PubMed Scopus Google the in Yates 3rd, J.R. A for protein identification and using tandem mass spectral data and protein sequence Chem. PubMed Scopus Google and D.L. Chambers M.C. accurate tandem mass spectral peptide identification by Proteome Res. 2007; PubMed Scopus Google and the in J. A. Mann M. a peptide search engine into the Proteome Res. PubMed Scopus Google Scholar). Although a from the of is these of calculated fragment ions. a using similar algorithms can from on how are for In many these tools the of the that this match by a in these tools a of match a of J. A. M. T. J. tandem mass spectrometry data PubMed Scopus Google to a the the of the match given that it the peptide for this spectrum. it the for the match given that the peptide is with the spectrum. The of the between these will be the match is more the the sequence is a on similar considerations J. T. M. fragmentation for interpretation of tandem mass spectrometry 5: PubMed Scopus Google Scholar). In Inc., a approach that to capture the and information of fragment ion is no of search that is for all of analysis all search The and on MS/MS search that in the are by the tools and analysis strategies that data search search tools and all a in is for parameter or search can be in and for With this is to considerations for common search and some that impact sequence database search performance for researchers new to the The first a user has in an MS/MS database search is sequence database to search MS/MS search tools a sequence database or from Scholar). Protein sequence databases of many are may a to search from a database many such as A. B. S. E. H. M. Martin C. The protein Res. 2005; PubMed Scopus Google Scholar). the sample of has these databases are protein sequence databases are available for the of searching acid databases or sequence or of these databases is the peptides in the sequence database may be by the search. a amino acid will the mass of the peptide to the sequence from with a spectrum. Researchers samples that may to protein databases that sequences. common to the sequence database is to identify the in the sample also to as many MS/MS spectra as proteins have been by into the proteins for the are and available protein samples proteins from the because of of these are in which sequences may be and the commonly for a The of database has become common search strategy for in protein identifications by mass 2007; PubMed Scopus Google Scholar). In the of a of protein sequences to be analysis to the to a for most researchers to of identification. A review will these in are of mass that are to a database peptide mass and fragment mass The peptide mass is on the mass accuracy of the used to measure the precursor ions masses in the MS scans. In various search this may be referred to as the peptide peptide mass or precursor The fragment mass is on the mass accuracy of the used to measure the fragment ions in the MS/MS scans. This parameter can be referred to as MS/MS fragment mass or product ion Although both are mass they are applied in a database search The peptide mass a in which peptides are with tandem mass spectrum. peptides the mass are subjected to fragment and The peptide mass parameter is on the mass accuracy of the MS1 spectra, for or of and for ion low mass accuracy the ion to for or see this parameter as a peptide the peptide mass for a a 5 to data the mass to will in the peptide sequence from to a tandem mass spectrum. In the mass to a will in search as more peptides are mass may in search because peptide match against a of with a to the is to that most search this parameter in mass that is to the or peptide instrument mass are usually measured in a for a precursor ion will to a in peptide a for a precursor ion will to a in peptide search engine of the precursor in the will to for the charge state in the measured mass and the peptide mass For the mass may be more a from the mass of the is for algorithms, or of instrument to the as the for a precursor the precursor mass to be by a mass To this a precursor to the mass many search an search that peptides from of mass This for an accurate mass search while also possible In this is using the in it is the mass and in the for the precursor ion search Fragment ion mass has been a for This is a parameter that how in MS/MS spectrum to search engine For a fragment to be an ion in the MS/MS spectrum be this of in For MS/MS scans in an ion a fragment of or in be with spectra from more accurate mass using are relatively to in the fragment because MS/MS fragmentation are specific to peptide fragment settings will so in using a fragment ion as will to is the most for liquid chromatography proteins with high peptides of amino acids in by the peptide to the of and is the of the is J. Pevzner P.A. trypsin Proteome Res. 2008; 7: PubMed Scopus Google Scholar). trypsin with some peptides that in the of or a typical enzyme search will trypsin and for or such as and are also used in proteomics these generate peptides that at different from sequence strategies are may be to or both of the peptide to to the of the A search peptides with a search peptides that have or or all possible peptides from a database. for or are in the for search and accurate mass and enzyme For will that are to the resulting in peptides or and this the of resulting peptides with to For et al. a between and the in which the was D.L. C. and specific trypsin of to of proteins in Chem. 2006; PubMed Scopus Google Scholar). fragmentation can in of MS/MS spectra of in the sample may be at is a of identifying possible or even no enzyme be The ability to identify peptides is a for complex in for modifications resulting from common sample In the most common of modification in MS/MS search algorithms, for a mass to or more amino acid are of modifications that a user can which have different on The first of modification is the mass of a given amino acid is to a different This modification is termed or One common modification that is applied in many proteomics is the addition of to to for modifications an amino acid mass impact on the database search in of and the search The of modification is a mass on a This modification is termed a or a modification and an of such a modification be the addition of to because of for or be common of a modification search. a search engine all of modified and in a modification the search of peptides is to search modifications the and be a peptide contains more capable of a such as a in a peptide with and the with which the modification can be to a is on the of fragment ions in the MS/MS spectrum from fragmentation between the possible of modification. search currently or no to the or of modification in In a identification by the Proteome of the of all of the a such as J. J. A approach for protein analysis and 2006; PubMed Scopus Google for the of for peptide spectrum is that some for modification will be into of most search into a search Cell. Full Full Scopus Google Scholar). for are protein modifications for mass PubMed Scopus Google and The database of protein modifications as a and PubMed Scopus Google Scholar). a sequence database search with a given tandem mass most that a peptide the spectrum or that no match to a spectrum can be in the database for a search a of peptides in a the is that all peptides are for This of peptides is because the for these peptides are used to the of the a identification can to the the peptide all of the peptides in the then that peptide is an of the and is a The in SEQUEST the peptides to a spectrum. is on the for all to the spectrum D. A method for the statistical of mass protein identifications using Chem. PubMed Scopus Google Scholar). For these and many search the of peptides and a key in the from given is a that that is similar to peptide or a separated from the various and are calculated on the of peptides it is to a of peptides to a in the calculated This is of using precursor mass that the of peptide with a enzyme search can in or a of peptides for spectral and has been shown to in identifications in proteomics analysis B. The with peptide and low Mascot scoring.J. Proteome Res. PubMed Scopus Google Scholar). This is to that all such are of identifications to of the are to of precursor mass accuracy in MS/MS data One method by many is to the peptide mass to a peptides that in the accurate mass The also by is to the accurate precursor mass analysis of the search This the search is using a or peptide mass are and to both accurate mass the search may to in the because of peptide search will be because the search is identify peptides for spectra that match a peptide a search a the search will and is to identifications because peptides are against many more a wide mass search to search by search also by mass identifications be at the accurate mass identifications be the mass and Mann to of mass including more mass on and mass in the database search Mann M. the of mass accuracy in Cell. Proteomics. 2007; Full Full PubMed Scopus Google Scholar). et al. that high mass accuracy search identification for low spectra, for peptides fragment ions are low in with acid and J. and of peptide mass accuracy in shotgun Cell. Proteomics. 2006; 5: Full Full PubMed Scopus Google Scholar). et al. Mascot and X!Tandem and that Mascot is more to in the peptide mass parameter with X!Tandem and benefit to accurate mass as a to a search M. Swamy S. Hubbard T. Choudhary J. Comparison of Mascot and X!Tandem performance for low and high accuracy mass spectrometry and the development of an adjusted Mascot threshold.Mol. Cell. Proteomics. 2008; 7: 962-970Abstract Full Full PubMed Scopus Google Scholar). et al. also that search accurate mass a of identifications at a given using SEQUEST B. Comparison of database search strategies for high precursor mass accuracy MS/MS Proteome Res. PubMed Scopus Google Scholar). against a precursor in to a search the enzyme Protein identification using MS/MS Proteomics. Scopus Google Scholar). a data analysis the has similar to the precursor mass The peptide search is in and so is a in of search with such analyses. with the in peptides for is a in an sequence the sequence for a spectrum. One to the enzyme is that peptides can be they even be in a search B. The of for shotgun Cell. Proteomics. 2007; Full Full PubMed Scopus Google Scholar). as the enzyme as a search analysis or as of tools L. J. for peptide identification from shotgun proteomics 2007; PubMed Scopus Google Scholar, M. L. Hubbard T. Choudhary J. and peptide identification with Mascot Proteome Res. PubMed Scopus Google and A. A.I. E. statistical to the accuracy of peptide identifications by MS/MS and database Chem. PubMed Scopus Google identify more peptides with using the enzyme the search P.A. for accurate protein identification in shotgun of and Biol. PubMed Scopus Google Scholar). is no on search strategies and analysis are user is to the of the identification of a low protein different analysis of an of analysis strategies to method for a For a shotgun of database to the of identifications at a is to perform this One of MS/MS analysis that has gained some is the ability to more a of proteins for or peptides with enzyme This is as it to be because of the search This of searching is implemented as in Mascot searching of tandem mass spectrometry PubMed Scopus Google in A method for the required to match protein sequences with tandem mass Mass PubMed Scopus Google mass with searching in and search. In a the algorithm in A. The a search engine that sequence and to identify peptides from tandem mass Cell. Proteomics. 2007; Full Full PubMed Scopus Google on proteins which sequence to more a search using a of sequence tags to and search This method the of data by the of database in such search searching spectra from the first against a of the of the can be by as using precursor against databases Pevzner approach and may Am. Soc. Mass Spectr. 22: PubMed Scopus Google Scholar). modification for a peptide may a in that it to be matched to a spectrum of an a are to be for modified database search and a or search in that the in the are all database to for the of a of in et al. discusses this and a with to database searching and C. statistical analysis for search Proteome Res. PubMed Scopus Google and a M. on statistical analysis for search Proteome Res. PubMed Scopus Google Scholar). mass spectrometry sequence database searching is an in the proteomics analysis Peptide identifications and protein in a or as are the product of most proteomics of the following will be in review they are The for identification D. A method for the statistical of mass protein identifications using Chem. PubMed Scopus Google Scholar, A. A.I. E. statistical to the accuracy of peptide identifications by MS/MS and database Chem. PubMed Scopus Google Scholar, of the SEQUEST Proteome Res. PubMed Scopus Google identifications from an L. J. for peptide identification from shotgun proteomics 2007; PubMed Scopus Google Scholar, D.L. Yates 3rd, J.R. and tools for and protein identifications from shotgun Proteome Res. PubMed Scopus Google Scholar, Z.Q. S. Chambers M.C. B. D.L. protein with high peptide identification Proteome Res. PubMed Scopus Google or with proteins and J. B. for proteins from peptide sequences inferred from tandem mass spectrometry Chem. 2007; PubMed Scopus Google Scholar, A.I. A. E. A statistical for identifying proteins by tandem mass Chem. PubMed Scopus Google Scholar, to protein from shotgun mass spectrometry Proteome Res. PubMed Scopus Google are key to the of these identifications from search to more the MS/MS data is commonly by tools such as Proteome a for PubMed Scopus Google and the L. D. T. H. E. B. B. Eng J.K. Martin D.B. A.I. A of the PubMed Scopus Google Scholar). A wide of tools more to the of a sample of these of and H. S. H. M. of and Res. PubMed Scopus Google A. H. B. A. A. A of protein and by Res. PubMed Scopus Google A. Wang D. B. T. a environment for of Res. PubMed Scopus Google and B. S. J. an for in various Res. 2005; PubMed Scopus Google Scholar). search is the of experimental tandem mass spectra to sequences. sequence databases and search identification and search While of may be through these algorithms, researchers may tools for data in which of modifications are E. M. F. identification of modified proteins using PubMed Scopus Google Scholar). search tools are most into with such as that perform and protein of peptides into protein and protein with an a of MS/MS scans may the field of identification will to
