A Scalable Approach for Protein False Discovery Rate Estimation in Large Proteomic Data Sets

Calculating the number of confidently identified proteins and estimating false discovery rate (FDR) is a challenge when analyzing very large proteomic data sets such as entire human proteomes. Biological and technical heterogeneity in proteomic experiments further add to the challenge and there are strong differences in opinion regarding the conceptual validity of a protein FDR and no consensus regarding the methodology for protein FDR determination. There are also limitations inherent to the widely used classic target–decoy strategy that particularly show when analyzing very large data sets and that lead to a strong over-representation of decoy identifications. In this study, we investigated the merits of the classic, as well as a novel target–decoy-based protein FDR estimation approach, taking advantage of a heterogeneous data collection comprised of ∼19,000 LC-MS/MS runs deposited in ProteomicsDB (https://www.proteomicsdb.org). The “picked” protein FDR approach treats target and decoy sequences of the same protein as a pair rather than as individual entities and chooses either the target or the decoy sequence depending on which receives the highest score. We investigated the performance of this approach in combination with q-value based peptide scoring to normalize sample-, instrument-, and search engine-specific differences. The “picked” target–decoy strategy performed best when protein scoring was based on the best peptide q-value for each protein yielding a stable number of true positive protein identifications over a wide range of q-value thresholds. We show that this simple and unbiased strategy eliminates a conceptual issue in the commonly used “classic” protein FDR approach that causes overprediction of false-positive protein identification in large data sets. The approach scales from small to very large data sets without losing performance, consistently increases the number of true-positive protein identifications and is readily implemented in proteomics analysis software. Calculating the number of confidently identified proteins and estimating false discovery rate (FDR) is a challenge when analyzing very large proteomic data sets such as entire human proteomes. Biological and technical heterogeneity in proteomic experiments further add to the challenge and there are strong differences in opinion regarding the conceptual validity of a protein FDR and no consensus regarding the methodology for protein FDR determination. There are also limitations inherent to the widely used classic target–decoy strategy that particularly show when analyzing very large data sets and that lead to a strong over-representation of decoy identifications. In this study, we investigated the merits of the classic, as well as a novel target–decoy-based protein FDR estimation approach, taking advantage of a heterogeneous data collection comprised of ∼19,000 LC-MS/MS runs deposited in ProteomicsDB (https://www.proteomicsdb.org). The “picked” protein FDR approach treats target and decoy sequences of the same protein as a pair rather than as individual entities and chooses either the target or the decoy sequence depending on which receives the highest score. We investigated the performance of this approach in combination with q-value based peptide scoring to normalize sample-, instrument-, and search engine-specific differences. The “picked” target–decoy strategy performed best when protein scoring was based on the best peptide q-value for each protein yielding a stable number of true positive protein identifications over a wide range of q-value thresholds. We show that this simple and unbiased strategy eliminates a conceptual issue in the commonly used “classic” protein FDR approach that causes overprediction of false-positive protein identification in large data sets. The approach scales from small to very large data sets without losing performance, consistently increases the number of true-positive protein identifications and is readily implemented in proteomics analysis software. Shotgun proteomics is the most popular approach for large-scale identification and quantification of proteins. The rapid evolution of high-end mass spectrometers in recent years (1.Scheltema R.A. Hauschild J.P. Lange O. Hornburg D. Denisov E. Damoc E. Kuehn A. Makarov A. Mann M. The Q Exactive hf, a benchtop mass spectrometer with a prefilter, high performance Quadrupole, and an ultra-high field Orbitrap analyzer.Mol. Cell. Proteomics. 2014; 13: 3698-3708Abstract Full Text Full Text PDF PubMed Scopus (232) Google Scholar, 2.Kelstrup C.D. Jersie-Christensen R.R. Batth T.S. Arrey T.N. Kuehn A. Kellmann M. Olsen J.V. Rapid and deep proteomes by faster sequencing on a benchtop Quadrupole ultra-high-field Orbitrap mass spectrometer.J. Proteome Res. 2014; 3: 6187-6195Crossref Scopus (137) Google Scholar, 3.Helm D. Vissers J.P. Hughes C.J. Hahne H. Ruprecht B. Pachl F. Grzyb A. Richardson K. Wildgoose J. Maier S.K. Marx H. Wilhelm M. Becher I. Lemeer S. Bantscheff M. Langridge J.I. Kuster B. Ion mobility tandem mass spectrometry enhances performance of bottom-up proteomics.Mol. Cell. Proteomics. 2014; 13: 3709-3715Abstract Full Text Full Text PDF PubMed Scopus (73) Google Scholar, 4.Yamana R. Iwasaki M. Wakabayashi M. Nakagawa M. Yamanaka S. Ishihama Y. Rapid and deep profiling of human induced pluripotent stem cell proteome by one-shot NanoLC-MS/MS analysis with meter-scale monolithic silica columns.J. Proteome Res. 2013; 12: 214-221Crossref PubMed Scopus (48) Google Scholar, 5.Hebert A.S. Richards A.L. Bailey D.J. Ulbrich A. Coughlin E.E. Westphall M.S. The Cell. Proteomics. 2014; 13: Full Text Full Text PDF PubMed Scopus Google proteomic that and as as proteins in a A. Hahne H. Wilhelm M. Kuster B. proteome analysis of the cell 2013; Full Text Full Text PDF PubMed Scopus Google Scholar, I. Mann M. to estimation in 2014; PubMed Scopus Google Scholar, M.S. K. K. M. strong for peptide of Proteome Res. 2013; 12: PubMed Scopus Google and of for the analysis of human and M. J. Hahne H. A. M. E. S. Marx H. Lemeer S. K. H. M. J. Bantscheff M. A. F. Kuster B. of the human 2014; PubMed Scopus Google Scholar, M.S. D. R. R. S. B. J. B. S. A. S. R. M. S.K. A. S. Y. A. S. J. R. S. K. A. J. D. M.S. S. A. H. C.J. S.K. R. A. T.S. H. A. of the human 2014; PubMed Scopus Google Scholar, H. D. D. R. S. Kuster B. Bantscheff M. Proteomics. in by profiling of the 2014; PubMed Scopus Google in most proteomic experiments is the identification of proteins in the proteins are by and tandem mass are used to protein sequence search that data to data in O. R. and of proteomic data by tandem mass PubMed Scopus Google Scholar, R. of proteomic the protein Cell. Proteomics. Full Text Full Text PDF PubMed Scopus Google are commonly by a search either a or a scoring A.L. approach to tandem mass data of with sequences in a protein PubMed Scopus Google Scholar, J. A. R.A. Olsen J.V. Mann M. a peptide search the Proteome Res. PubMed Scopus Google Scholar, D.J. protein identification by sequence mass spectrometry PubMed Scopus Google Scholar, R. proteins with tandem mass PubMed Scopus Google Scholar, M. mass spectrometry search Proteome Res. 3: PubMed Scopus Google are from identified and a protein or a as a for the in the identification R. of proteomic the protein Cell. Proteomics. Full Text Full Text PDF PubMed Scopus Google Scholar, O. of for protein identification tandem mass PubMed Google the of false discovery in an is to and the of protein identifications. to conceptual and the most widely used strategy to FDR in proteomics is the target–decoy search strategy search strategy for in large-scale protein identifications by mass PubMed Scopus Google The this is that with in the target and the decoy or of the same K. S. discovery in Scopus Google Scholar, of and rate estimation for peptide and protein identification in Proteomics. PubMed Scopus Google The number of to the decoy an of the number of to in the target The number of target and decoy used to either a or a FDR for a data K. S. discovery in Scopus Google Scholar, of and rate estimation for peptide and protein identification in Proteomics. PubMed Scopus Google Scholar, H. discovery and in mass Proteome Res. PubMed Scopus Google Scholar, and false discovery of the same Proteome Res. PubMed Scopus Google Scholar, of novel decoy for protein identification data Proteome Res. PubMed Scopus Google Scholar, S. for false and false discovery in PubMed Scopus Google to the FDR the of and by the number of target and decoy a score. the of the that a peptide FDR that in an protein FDR for a small or data an high protein FDR when the data of and rate estimation for peptide and protein identification in Proteomics. PubMed Scopus Google Scholar, M. M. A. R. identification false discovery for very large proteomics data sets by tandem mass Cell. Proteomics. Full Text Full Text PDF PubMed Scopus Google is the of the is when a large of the true positive proteins In small data a to a peptide over decoy and target for a estimation of false positive target identifications by the number of decoy identifications. in large experiments to of LC-MS/MS or target proteins and an number of proteins to by false positive peptide In decoy proteins are by the peptide the number of false positive protein identifications from the decoy The the number of identified target proteins the this this is for in the decoy an of false and and a for for the of false positive protein M. M. A. R. identification false discovery for very large proteomics data sets by tandem mass Cell. Proteomics. Full Text Full Text PDF PubMed Scopus Google the that protein identifications false positive are over the target the number of false positive protein identifications a are from the number of protein and the number of target and decoy protein identifications. The protein FDR is by the number of false positive identifications of the by the number of target identifications. this approach was for large data sets on LC-MS/MS runs from of is the approach strategy for of false positive the was for D. H. R. analysis of proteomic data peptide and protein identification and Cell. Proteomics. Full Text Full Text PDF PubMed Scopus Google and for proteins of and protein for the of protein identification by tandem Proteome Res. 2014; 13: PubMed Scopus Google of and decoy in the range is the number of true peptide or protein identifications is to to and The number of decoy is by the when FDR The approach is than the strategy and to is also based on the that the of the decoy in the classic target–decoy strategy to the same in In the of the is to that there is no consensus in the regarding and protein for data of is to and to the peptide to protein and of decoy protein protein J. protein FDR 2013; Google is as an of protein in proteomic experiments is that target–decoy are when of individual are approach and false discovery when PubMed Scopus Google to the or peptide from analysis to to a protein FDR D. of the human proteome in as the and proteomes for the and Proteome Proteome Res. 2014; 13: PubMed Scopus Google is We an protein FDR approach “picked” target–decoy strategy that performance over the in a very large proteomic data M. J. Hahne H. A. M. E. S. Marx H. Lemeer S. K. H. M. J. Bantscheff M. A. F. Kuster B. of the human 2014; PubMed Scopus Google a of the performed the In this study, we further the for protein FDR estimation and investigated with that of the classic FDR in data sets of to ∼19,000 LC-MS/MS The show that the is in decoy protein true positive and well for small and large proteomic data sets. The data for this was a large collection of LC-MS/MS runs with the human protein identification data deposited in ProteomicsDB (https://www.proteomicsdb.org). the of this comprised LC-MS/MS the of which of the human proteome M. J. Hahne H. A. M. E. S. Marx H. Lemeer S. K. H. M. J. Bantscheff M. A. F. Kuster B. of the human 2014; PubMed Scopus Google Scholar, M.S. D. R. R. S. B. J. B. S. A. S. R. M. S.K. A. S. Y. A. S. J. R. S. K. A. J. D. M.S. S. A. H. C.J. S.K. R. A. T.S. H. A. of the human 2014; PubMed Scopus Google In are experiments of number of LC-MS/MS from each in D.J. protein identification by sequence mass spectrometry PubMed Scopus Google and J. A. R.A. Olsen J.V. Mann M. a peptide search the Proteome Res. PubMed Scopus Google Scholar, J. Mann M. high peptide identification mass and protein PubMed Scopus Google a protein sequence the human proteome and of as M. J. Hahne H. A. M. E. S. Marx H. Lemeer S. K. H. M. J. Bantscheff M. A. F. Kuster B. of the human 2014; PubMed Scopus Google in the and The with the target–decoy search a decoy with protein a of and a of for and for an of or a of the of and of as well as of protein as and as for individual experiments stable with in cell tandem mass or In the the same target–decoy protein sequence as the search and as of the was by the of and mass was to and for and regarding and data in M. J. Hahne H. A. M. E. S. Marx H. Lemeer S. K. H. M. J. Bantscheff M. A. F. Kuster B. of the human 2014; PubMed Scopus Google data to the in this as well as the protein are in and search for data sets in are in engine-specific peptide as in Wilhelm M. J. Hahne H. A. M. E. S. Marx H. Lemeer S. K. H. M. J. Bantscheff M. A. F. Kuster B. of the human 2014; PubMed Scopus Google as peptide of the same for and in of and by a with a of to for by the scoring The false discovery in each by the number of decoy by the number of target and the was a with a of to for small The over with a false discovery rate than was to the peptide of by the or by the peptide the of this study, a q-value is to the FDR which a or protein in the are commonly used to a of to a of for this we the with the highest search that peptide sequence and a peptide for each LC-MS/MS in by or by the from to and the number of by the number of a from to the q-value from the to to the q-value the and was by a the highest and scoring with an q-value as the and of the q-value a by the with the a and the the of was to either protein or to protein from the same are as are as from protein the of this study, is the identification of a protein and the identification of protein of a data in protein as the of of the best scoring peptide protein either as the of the of that a q-value or by the of protein proteins in by score. protein by the from to and the number of by the number of a from to the q-value to the q-value was each a data was we with the the number of identifications by the with the number of and was to that the number of protein FDR a and data in data in In to the classic the treats target and decoy sequences of the same protein as a the protein for the target sequence is than that of the decoy the target sequence is as a and the decoy sequence is the decoy sequence than the target as a decoy and the target sequence is no is with to target and decoy proteins to the protein The protein FDR was the target and decoy in the same as in the classic In large proteomic or of of the classic target–decoy strategy protein FDR the the number of identified target proteins the the of target and decoy protein identifications to an of decoy proteins and protein this we used protein identification from a of LC-MS/MS and protein identification protein FDR the classic of each LC-MS/MS FDR and search in to the number of proteins protein by of the best for of that on the search target proteins and decoy proteins with identified proteins protein We the search and on and the protein FDR estimation each that the number of identified target proteins when further search and that by the search protein identifications a rate the number of target as the number of search an target proteins identified when the search the protein FDR search the target protein by proteins also the classic FDR to that of proteins is that the the search proteins The by a protein FDR a protein FDR each protein the search to proteins when further search this of the classic protein FDR approach, we to an we to as the “picked” target decoy strategy this the heterogeneous of the data in ProteomicsDB data that is in the The human proteome data deposited in ProteomicsDB from a wide of and experiments and was on of Orbitrap and as well as the data to and in a that a and of the the of ProteomicsDB Orbitrap for which was used as a search and for which was We differences in the of the search which is in the differences in the scoring In we and a in and for to peptide for we for and M. J. Hahne H. A. M. E. S. Marx H. Lemeer S. K. H. M. J. Bantscheff M. A. F. Kuster B. of the human 2014; PubMed Scopus Google Scholar, K. S. discovery in Scopus Google Scholar, J. Mann M. high peptide identification mass and protein PubMed Scopus Google Scholar, D. of search in proteomics.Mol. Cell. Proteomics. 2013; 12: Full Text Full Text PDF PubMed Scopus Google We also that the target decoy are on the of and the of used or of human stem by very target–decoy with of the cell by is to a to FDR in heterogeneous and large data sets. for each LC-MS/MS is by or implemented in J. Mann M. high peptide identification mass and protein PubMed Scopus Google for and and false discovery of the same Proteome Res. PubMed Scopus Google or A. E. R. to the of peptide identifications by and PubMed Scopus Google for In to for and data we implemented a simple for q-value with search of for this we the highest scoring that peptide sequence that and is with a the best to we and that the of the the same the best or the best peptide K. S. discovery in Scopus Google Scholar, the of estimation for in Proteomics. 2013; PubMed Scopus Google in that are by the of high with A. Bantscheff M. of data analysis for mass PubMed Scopus Google we to the and to in to with the that the number of decoy is very small for high scoring there no in q-value a of of and the with the than the scoring peptide for that to from the search the of to as very well particularly most of the false are and decoy that the q-value of a as a of the of the data and lead to a in FDR the data in we investigated of false positive protein identifications with the classic The q-value was to protein we used the and that we used best scoring peptide for protein peptide with the best for protein of and rate estimation for peptide and protein identification in Proteomics. PubMed Scopus Google Scholar, D. H. R. analysis of proteomic data peptide and protein identification and Cell. Proteomics. Full Text Full Text PDF PubMed Scopus Google the best peptide or the of peptide are in the proteomics The of target and decoy proteins to the is in the that the range false positive protein identifications A. E. R. to the of peptide identifications by and PubMed Scopus Google the same the number of decoy proteins in that range is than that of the target proteins the of false positive proteins. We investigated an approach that we “picked” In to the classic the treats target and decoy sequences of the same protein as a pair rather than as individual the protein for the target sequence is than that of the decoy the target sequence is as a and the decoy sequence is the decoy sequence than the target as a decoy and the target protein is was in by the decoy approach used for J. B. M. D. B. sequencing search for and peptide Cell. Proteomics. Full Text Full Text PDF Scopus Google and in by the in the field of a target and decoy in to for that the best in either the target or the decoy rather than that a in target and decoy search strategy for in large-scale protein identifications by mass PubMed Scopus Google Scholar, the of estimation for in Proteomics. 2013; PubMed Scopus Google or approach is as no for the of either a decoy protein or a false positive target the the target is the decoy is to the range of the target as for FDR approach A. E. R. to the of peptide identifications by and PubMed Scopus Google a the of true positive protein identifications the the target and decoy in for very protein that the estimation of false positive is also to the approach of and protein for the of protein identification by tandem Proteome Res. 2014; 13: PubMed Scopus Google that of decoy by an The decoy positive and for true positive protein identification protein a than the classic approach when the of of of a protein as a we a the of false positive and true positive protein which to the that large decoy proteins high by of scoring We the classic and for to true positive proteins in the when the differences target and decoy protein identifications q-value very of proteins for q-value the number of true positive identifications for the classic the number of true positive protein identifications for the a stable proteins. The true positive as a of q-value is a of a FDR estimation A. E. R. to the of peptide identifications by and PubMed Scopus Google protein FDR in the same the classic protein FDR for of and the protein FDR a of the protein FDR performance that the is a and protein FDR estimation We the analysis in and the number of identified target and decoy proteins as a of and we a q-value of used the best for protein scoring as the experiments in a and and search The data was the classic as well as the is that the target proteins than the classic a number of decoy protein identifications and We that the faster than the when the classic the the decoy protein show the an very the number of that of data the that a protein as a false positive identified is by a high in the The are in the protein FDR and the protein FDR increases for the classic as the data for the we the data protein the number of confidently identified proteins increases for the classic and the as the data the is consistently and the of identified proteins also increases as the data In the data the classic approach proteins protein FDR proteins are with the is that the approach for this We the data analysis strategy to the of data in to on a mass spectrometry based of the human proteome M. J. Hahne H. A. M. E. S. Marx H. Lemeer S. K. H. M. J. Bantscheff M. A. F. Kuster B. of the human 2014; PubMed Scopus Google the classic FDR strategy proteins protein FDR with proteins the the strategy without protein proteins of the target protein FDR to true positive protein identifications in the data analyzing the of the data of the proteome M.S. D. R. R. S. B. J. B. S. A. S. R. M. S.K. A. S. Y. A. S. J. R. S. K. A. J. D. M.S. S. A. H. C.J. S.K. R. A. T.S. H. A. of the human 2014; PubMed Scopus Google and a number of further data the number of protein identifications FDR to and proteins the strategy to this data without protein proteins of the target protein FDR to true positive protein identifications in the data in the analysis is the that the best for a protein is very with to which q-value is the of protein identification the of of for a protein are to an q-value and high The the of as well as the best approach a FDR of is to that FDR FDR lead to of false peptide identifications in the data and of data analysis such as identification of and protein and in FDR also a of peptide data are the for each LC-MS/MS in the data to and protein FDR FDR the classic and FDR with the and a peptide and In the we that the the for very large data sets. for small or individual In no the in protein identifications the in protein as the number of protein identifications in a we the to the of a number of large-scale protein identification M. J. Hahne H. A. M. E. S. Marx H. Lemeer S. K. H. M. J. Bantscheff M. A. F. Kuster B. of the human 2014; PubMed Scopus Google Scholar, J. Mann M. of human novel Res. PubMed Scopus Google Scholar, A. J. Mann M. proteomic analysis of cell of most Cell. Proteomics. Full Text Full Text PDF Scopus Google Scholar, J. A. A. The proteomes of pluripotent stem and stem PubMed Scopus Google Scholar, J. M. J. S. Mann M. proteome and of a human cell PubMed Scopus Google Scholar, J. C.D. S. Bailey D.J. R. and of human and PubMed Scopus Google Scholar, J. S. S. of proteins in the and in the as of the Proteome Proteome Res. 2013; 12: PubMed Scopus Google Scholar, J. proteomic analysis of by 2013; PubMed Scopus Google and that the consistently identified a number of proteins than the classic We that differences in data the search and used and also to the differences to protein identifications In this study, we investigated the and performance of the “picked” target decoy strategy for estimating protein false discovery in large proteomics data sets. The decoy protein for the classic and that the of a false positive is for proteins. large target and decoy proteins are to high scoring and are to protein than small proteins of which the protein that to or are the number of the of the number of of mass spectrometer and used and of by simple data and the of from the commonly approach of target and decoy sequences for to target and decoy of a protein sequence as a proteins that in target and decoy the with the highest and the this approach the of decoy for the classic FDR the target protein The of target and decoy in the or no the performance of the in with on the of the A. E. R. to the of peptide identifications by and PubMed Scopus Google The also to the approach that for over-representation of decoy by the with an of and protein for the of protein identification by tandem Proteome Res. 2014; 13: PubMed Scopus Google The performance of the is a in target decoy the approach for this a simple of decoy is the regarding or a decoy peptide is in a decoy this is by peptide sequences to the target that is used for protein identification is to the that a decoy sequence a of a peptide sequence or a or this number of each high scoring decoy protein identification and the protein there are such the number of proteins that identified in a proteome the of protein FDR a a that no the mass data The analysis further that protein scoring the best for a protein performed than for a is the is to protein and that the of peptide in large data sets a on protein scoring the number of the search or lead to the approach, which is to scoring of and rate estimation for peptide and protein identification in Proteomics. PubMed Scopus Google peptide of protein this the of protein and peptide protein scoring or of and in to the classic FDR that protein FDR as the data the number of decoy is data when the is an a false positive protein identification by a scoring target or decoy to a true high scoring target when high a tandem is to the data is that data to an large data add false is a as as proteome identification is the of the data a novel protein or identified In on a human M. J. Hahne H. A. M. E. S. Marx H. Lemeer S. K. H. M. J. Bantscheff M. A. F. Kuster B. of the human 2014; PubMed Scopus Google we the of the protein for of the protein in the we in We further that we to a estimation of protein FDR for such a very large data that The we as a of the FDR approach a of the number of identified proteins in that protein identifications in the same and to to of the we implemented a that the identifications in and depending on the from this analysis is that in to confidently proteins in a proteome by large of high LC-MS/MS data that the protein of an with the that the FDR approach performed consistently than the classic FDR for of we that this approach is and in proteomic software. We for with the with discovery rate induced scoring peptide combination q-value

A Scalable Approach for Protein False Discovery Rate Estimation in Large Proteomic Data Sets | Litlas