Multibatch TMT Reveals False Positives, Batch Effects and Missing Values
Multiplexing strategies for large-scale proteomic analyses have become increasingly prevalent, tandem mass tags (TMT) in particular. Here we used a large iPSC proteomic experiment with twenty-four 10-plex TMT batches to evaluate the effect of integrating multiple TMT batches within a single analysis. We identified a significant inflation rate of protein missing values as multiple batches are integrated and show that this pattern is aggravated at the peptide level. We also show that without normalization strategies to address the batch effects, the high precision of quantitation within a single multiplexed TMT batch is not reproduced when data from multiple TMT batches are integrated.Further, the incidence of false positives was studied by using Y chromosome peptides as an internal control. The iPSC lines quantified in this data set were derived from both male and female donors, hence the peptides mapped to the Y chromosome should be absent from female lines. Nonetheless, these Y chromosome-specific peptides were consistently detected in the female channels of all TMT batches. We then used the same Y chromosome specific peptides to quantify the level of ion coisolation as well as the effect of primary and secondary reporter ion interference. These results were used to propose solutions to mitigate the limitations of multi-batch TMT analyses. We confirm that including a common reference line in every batch increases precision by facilitating normalization across the batches and we propose experimental designs that minimize the effect of cross population reporter ion interference. Multiplexing strategies for large-scale proteomic analyses have become increasingly prevalent, tandem mass tags (TMT) in particular. Here we used a large iPSC proteomic experiment with twenty-four 10-plex TMT batches to evaluate the effect of integrating multiple TMT batches within a single analysis. We identified a significant inflation rate of protein missing values as multiple batches are integrated and show that this pattern is aggravated at the peptide level. We also show that without normalization strategies to address the batch effects, the high precision of quantitation within a single multiplexed TMT batch is not reproduced when data from multiple TMT batches are integrated. Further, the incidence of false positives was studied by using Y chromosome peptides as an internal control. The iPSC lines quantified in this data set were derived from both male and female donors, hence the peptides mapped to the Y chromosome should be absent from female lines. Nonetheless, these Y chromosome-specific peptides were consistently detected in the female channels of all TMT batches. We then used the same Y chromosome specific peptides to quantify the level of ion coisolation as well as the effect of primary and secondary reporter ion interference. These results were used to propose solutions to mitigate the limitations of multi-batch TMT analyses. We confirm that including a common reference line in every batch increases precision by facilitating normalization across the batches and we propose experimental designs that minimize the effect of cross population reporter ion interference. Highlights•Revealed inflation of missing values as multiple TMT 10-plex batches are integrated.•Analyzed the impact of integrating multiple TMT 10-plex batches on the quantification accuracy of both high and low abundance proteins.•Established reliable detection of false positives caused by coisolation and reporter ion interference, highlighted by the incidence of Y chromosome peptides in all female channels.•Optimized new experimental design set-ups to minimize cross population reporter ion interference via insights into coisolation and reporter ion interference. High-throughput, shotgun proteomics, using data dependent acquisition (DDA), 1The abbreviations used are:DDAdata dependent acquisitionPTMpost-translational modificationRIIreporter ion interferenceCIIcoisolation interferenceTMTtandem mass tags. 1The abbreviations used are:DDAdata dependent acquisitionPTMpost-translational modificationRIIreporter ion interferenceCIIcoisolation interferenceTMTtandem mass tags. now enables the comprehensive study of proteomes, allowing the identification of 10,000 or more proteins from cells and tissues (1Bekker-Jensen D.B. Kelstrup C.D. Batth T.S. Larsen S.C. Haldrup C. Bramsen J.B. Sorensen K.D. Hoyer S. Orntoft T.F. Andersen C.L. Nielsen M.L. Olsen J.V. An Optimized Shotgun Strategy for the Rapid Generation of Comprehensive Human Proteomes.Cell Syst. 2017; 4: 587-599 e584Abstract Full Text Full Text PDF PubMed Scopus (255) Google Scholar, 2Beck M. Schmidt A. Malmstroem J. Claassen M. Ori A. Szymborska A. Herzog F. Rinner O. Ellenberg J. Aebersold R. The quantitative proteome of a human cell line.Mol. Syst. Biol. 2011; 7: 549Crossref PubMed Scopus (585) Google Scholar, 3Meier F. Geyer P.E. Virreira Winter S. Cox J. Mann M. BoxCar acquisition method enables single-shot proteomics at a depth of 10,000 proteins in 100 minutes.Nat. Methods. 2018; 15: 440-448Crossref PubMed Scopus (218) Google Scholar). to proteome using of mass is (1Bekker-Jensen D.B. Kelstrup C.D. Batth T.S. Larsen S.C. Haldrup C. Bramsen J.B. Sorensen K.D. Hoyer S. Orntoft T.F. Andersen C.L. Nielsen M.L. Olsen J.V. An Optimized Shotgun Strategy for the Rapid Generation of Comprehensive Human Proteomes.Cell Syst. 2017; 4: 587-599 e584Abstract Full Text Full Text PDF PubMed Scopus (255) Google Scholar, S. The of protein and peptide mass in A. PubMed Scopus Google Scholar). evaluate the of the a of for is also Aebersold R. quantitative data for Biol. PubMed Scopus Google Scholar, of The of protein Full Text Full Text PDF PubMed Scopus Google Scholar). The data acquisition is for that the of the for in protein and M. proteomics for cell Biol. PubMed Scopus Google Scholar, M. F. M. quantitative proteomics of proteome in human Full Text Full Text PDF Scopus Google Scholar, M. M. A. protein using in and mass protein 15: Full Text Full Text PDF PubMed Scopus Google Scholar). data dependent acquisition reporter ion interference coisolation interference tandem mass tags. data dependent acquisition reporter ion interference coisolation interference tandem mass tags. with the of large-scale proteomics strategies have to multiple to be in peptides M.L. S. F. J. J.B. A. proteome analyses of human of 2018; PubMed Scopus Google Scholar, J. F. M. R. M. J. of the J. 2018; PubMed Scopus Google Scholar). The used TMT A. J. S. J. Schmidt R. C. mass a quantification for of protein by PubMed Scopus Google and S. S. S. S. S. S. M. F. A. protein quantitation in using Full Text Full Text PDF PubMed Scopus Google tags for peptide identification and TMT in and is now used M. S. J. protein PubMed Scopus Google Scholar, M. R. enables and multiplexed detection of across cell line PubMed Scopus Google Scholar). the of multiplexed TMT to in proteomics and the that from the in proteomics C. M. C. for the multiple of missing values in quantitative proteomics data to 15: PubMed Scopus Google Scholar, J. K.D. and of the of missing for mass PubMed Scopus Google Scholar). within a single multiplexed TMT the of missing values at the protein level is M. S. J. protein PubMed Scopus Google Scholar). Further, the precision of the quantification within a multiplexed TMT batch is high of common protein quantification 2018; PubMed Scopus Google Scholar). is well multiplexed TMT for large-scale TMT batches. this we a proteomic data set of human iPSC 10-plex TMT batches A. The iPSC proteomic 2018; Scholar). We the quantitation of data both within and 10-plex batches and on missing accuracy of false positives and the effect of both reporter ion interference and coisolation interference We show is an effect on missing values as data from multiple batches are integrated both at the protein and peptide level. We both by the of within 10-plex TMT and by a reference line of the iPSC line that were common to every the incidence of false positives was studied by using Y chromosome peptides as an internal control. The iPSC lines quantified in this data set were derived from including both male and hence the peptides mapped to the Y chromosome should be absent from female lines. Nonetheless, we confirm that these Y chromosome-specific peptides were consistently detected in the female channels of all TMT batches. by using these Y chromosome we quantified the effect of ion coisolation and reporter ion interference TMT quantification The study of iPSC and derived from The study twenty-four 10-plex TMT batches. batch of common reference line of iPSC cell line and iPSC cell lines. The were used for the data normalization of the were derived from female and from male protein iPSC cell were with and in of in 100 and at for was using on The proteins were using for at then in the for using protein was quantified using the the with mass the were with 100 then a with and were used at an to of The were at then by with to a of were using tandem mass the peptides were in 100 and was using a 10-plex TMT batch 100 of peptides from cell line to be in 100 of were with a TMT in for at the was using of for and the cell were and in The TMT were using were a with a the were using a of at and in at a rate of were into were into The were and the peptides in and by were using an mass with a was using a were a and on a integrated using a from to with a rate of The and The was by to the and the data were the of in a using and The is in the the from to with a mass of and an of The were for using in the ion with and an of The was set to with a of and a of the rate was set to the for more TMT were using with a of and using of The were then in the with a of The was set to and the was set to of the TMT batches were on the same TMT was by of a peptide to evaluate was by of an cell to evaluate peptide and protein The of of with an and with the the to be The data from all twenty-four 10-plex TMT batches were batches were using J. Mann M. enables high peptide identification mass and protein PubMed Scopus Google Scholar, S. Cox J. The for mass shotgun PubMed Scopus Google The was set to for of the and The data was with the was set to ion with of to with a reporter mass set to peptide was set to and peptides were identified using The are at R. A. F. M. A. S. R. proteomics data and PubMed Scopus Google via the A. J. F. R. of the and PubMed Scopus Google with the J. Mann M. enables high peptide identification mass and protein PubMed Scopus Google quantification proteins that were as or identified by were The as or were also from the analysis. The peptide data set were the proteomic Cox J. Mann M. for protein and without Full Text Full Text PDF PubMed Scopus Google Scholar). is the protein is is the mass of the protein protein is the protein and is the for all These be to as were used to study the of for the 10-plex a was to every in every to the protein is the protein derived from reporter reference The is for in all and all channels to missing values within this a of that were detected with at reporter were for the of missing values within 10-plex TMT the of reporter was with the of identified within the was to the missing for of the 10-plex TMT batches. the effect of integrating multiple TMT was to missing values are by a in the of 10-plex TMT batches was in an from and with batches was not used for this with level. the level batches be at with and at the level batches be at with was with the of the level a new of detected with at reporter ion within of the integrated TMT batches was and the of with reporter was the new The of in protein abundance was using the protein protein the is to the by the The protein within 10-plex TMT batch was for all cell lines within the same using all proteins detected in every reporter The reference line was using proteins that were detected in the across all of the 10-plex TMT batches. 10-plex TMT a was for all cell lines within the same The were using from the The same was to the values for the reference using reporter in all TMT batches. The was The for is the of all and The is the of for all The reporter ion interference are on a data for 10-plex TMT from as in ion interference for all TMT the reporter mass the reporter within the and the channels for primary and secondary reporter ion in a new study the effect of reporter ion interference across TMT we a of peptides that were specific to the of protein on the Y and of using peptide values from Y chromosome specific a of male and female iPSC lines in 10-plex TMT of the TMT batches female not to have Y chromosome derived in analyses A. A. S. S. A. A. R. F. Mann A. A. R. S. M. A. R. A. R. O. in human 2017; PubMed Scopus Google Scholar). these female peptide to Y chromosome specific was from the analysis. An an and was hence also from the analysis. of Y chromosome-specific peptides were used for this data for The peptide male channels female to reporter ion interference were 10-plex TMT for using the The male to the reporter ion interference used these peptide batch and was using for Google Scholar). The peptide reporter ion interference in female channels to female channels with reporter ion interference, were within 10-plex TMT for using the These results were by the peptides with or to the were and the were The reporter ion interference to the not by reporter ion interference used these peptide batch and was using for Google Scholar). of using TMT is the low of missing values that are within a single TMT as low as missing values at the protein level of common protein quantification 2018; PubMed Scopus Google data are not at the peptide level. We by the iPSC 10-plex TMT data for the of missing values at the protein level within TMT batch The results are with of the 10-plex TMT batches show missing values at the protein with with missing protein values was experiment is highlighted in and Further, when we the data at the peptide is to the protein with of the 10-plex TMT batches missing peptide the batch an effect and missing values at the peptide level. We from the of the analysis. These results not address the effect of integrating data from 10-plex TMT batches into a single analysis. study the effect of data we the of batches from to and the of missing values that were and the protein the of missing values increases from with 10-plex TMT to when data from a 10-plex TMT batch were integrated we data from 10-plex TMT the of missing values at the protein level to was when the was at the peptide level integrating data from 10-plex TMT the of missing peptide values was more integrating data from 10-plex TMT batches to missing values at the peptide level. The data peptides are not detected is low abundance these we to a more on the inflation rate of peptide missing We that the of peptides identified within 10-plex TMT batch is across batches. The of peptides identified batch was with a of these peptide level we the for all peptides in all cell lines The of We the peptide data set by on the values The the peptides and the the are peptides within the that are detected in all TMT channels and peptides that are in of the TMT the the a A. A. in new in Scopus Google is this is when we the results from the we that of these peptides are detected in of the TMT These high peptides that are detected in of all channels have a of the of of the peptides from the were detected in all the TMT are peptides that were detected in all TMT of of all we the data by the identification by the of TMT channels in were detected The the of peptides that were detected in TMT channels these peptides a the that peptides are not identified of peptides are detected in of all TMT have TMT as a method in a of data of common protein quantification 2018; PubMed Scopus Google Scholar). of these have on quantitative precision within a single TMT and not the effect of integrating data from multiple TMT batches into analysis. large proteomic analyses of multiple cell lines to multiple TMT batches in a single experiment M.L. S. F. J. J.B. A. proteome analyses of human of 2018; PubMed Scopus Google Scholar). We protein for iPSC and of a iPSC line across 10-plex TMT batches is from the We then to the and is to evaluate to evaluate PubMed Scopus Google Scholar). was for every iPSC line within TMT 10-plex and for all the of the across the 10-plex TMT batches. The within 10-plex TMT batch is the precision of the quantitation within single when the same is to the of the iPSC line across the to this we the for the protein M. across the and Scopus Google both within 10-plex TMT and across we the protein within 10-plex TMT the was with all a protein the data show that for every proteins with a were These data show high precision of quantitation within multiplexed to evaluate accuracy we the for all of the reference cell were in in every 10-plex TMT The of all the proteins detected in the was that the The of of all proteins in the be in all the 10-plex TMT analyses. is in proteomics that low proteins and We to this and on the We the on the and the 100 proteins the was and on across all reference line data We to protein highlighted in highlighted in highlighted in highlighted in and highlighted in in from to and was identified with peptides from to and was identified with from to and was identified with from to and was identified with and from to and was identified with of these proteins are with a within the of the reference have identified with and the multiplexed that the is not to low abundance We also that the of iPSC lines in this study from donors, of these TMT for of iPSC lines derived from both and with including and Nonetheless, the protein within is the from the of that TMT batch have a on the proteomics data a results that TMT is a and for quantitative proteomics, is to be also of when data from multiple TMT batches. is to that a of normalization Cox J. Mann M. for protein and without Full Text Full Text PDF PubMed Scopus Google to with batch These that when large-scale proteomics analyses across multiple TMT is to be of the for batch to data the batch have from protein S. a for protein using PubMed Scopus Google to a reference line from multiple without using common reference PubMed Scopus Google Scholar). Here we used the of a reference iPSC line as an internal reference to for batches this normalization method a of across all cell lines and the results to the for analysis. The iPSC data set A. The iPSC proteomic 2018; Scholar, A. A. S. S. A. A. R. F. Mann A. A. R. S. M. A. R. A. R. O. in human 2017; PubMed Scopus Google with an to study the incidence of false positives within a TMT The study used iPSC lines derived from both male and female within of the twenty-four 10-plex TMT batches the lines from male should proteins by on the Y this a set of we to the of false positives as well as the of reporter ion interference TMT channels and coisolation interference The data set detected proteins that were mapped to the Y all peptides derived from these proteins should be in the TMT channels with male cell lines in should be absent in the TMT channels with female cell lines. from we on a of peptides that mapped to the Y chromosome specific of the 10-plex TMT batches and female cell Y chromosome-specific peptides that were detected in these batches were as and from analysis. batch was also an and not for analysis. a we on Y chromosome peptides that were used as for the of false positives within the TMT batches. We false positives by the female TMT channels were from Y chromosome-specific peptides this that in all 10-plex TMT batches and in all reporter channels a female cell a of of the Y chromosome-specific peptides identified within the batch also in the female across all these a of of Y chromosome-specific peptides quantified in batch were quantified in TMT channels that female cell lines. The Y chromosome peptides should not be in female hence the level of detection is We that the of for Y chromosome-specific peptides in the channels female cell lines results from a of coisolation and reporter ion interference. ion interference, also as from level and experimental M. J. C. in and the and the PubMed Scopus Google Scholar). interference is the effect caused by multiple peptides within the J. J. R. O. Ori A. of protein quantification in a by and TMT with PubMed Scopus Google Scholar). study both we on the Y chromosome-specific as these should be in the male channels and absent in the female detected in female channels should be We used the Y chromosome peptides to evaluate the in peptide male and female across all the 10-plex TMT batches The results of the significant TMT batches. as and have and in male and female the detection of false positives of coisolation interference. and show a and the detection of the false positives We both batches low peptides and hence low female channels more to coisolation interference. evaluate coisolation ion interference, we female channels with primary or secondary reporter ion interference as of coisolation interference for ion interference in PubMed Scopus Google Scholar). in male channels show a of with female channels not by reporter ion interference. the on the peptide the peptide across male lines was or to the the peptide of male lines was the to coisolation interference. We also the of reporter ion interference this we a peptide specific for the male channels the of reporter ion interference in female when a male by the female of the same to female when a male by the female of the same to female when a female is by both and from male not by or secondary were as The to the were used to the The male lines were a of female channels not by reporter ion interference female channels to primary and secondary reporter ion interference The effect was caused by the male lines were the female that is the of and we the and We that for both and the false positives are within the for in detected within proteomic data The cell proteome and by the PubMed Scopus Google Scholar, A. of the to cell in human Scopus Google Scholar). quantify the primary and secondary reporter ion interference across peptide abundance we now female channels by reporter ion interference and to female channels not by reporter ion interference We also this by the peptide with high values or the and low the the we that the effect was caused by for high Y chromosome-specific peptides a with the channels not by reporter ion interference and in low abundance peptides The a more with a of in the high peptides and in the low The of primary and secondary a of in the high peptides and a in the for the low These results low peptides are by coisolation interference, as reporter ion interference to effect on that reporter ion interference have effect in quantification of high that the design of TMT experiment to minimize the on data quantification of cross reporter ion interference. all on more a single TMT we at internal reference should be in batch and to or These channels the of and are by a of in the reference line at or increases the impact of reporter ion interference by to results also show TMT experimental designs that to minimize the of primary and secondary reporter ion interference the in a 10-plex TMT when are with a multiple channels to be by cross reporter ion interference The design the across the channels a is or for in or a and we using TMT as all 10-plex TMT cross reporter ion interference. An TMT set enables a design without reporter ion interference the channels at and to this a is as then should be in the to the experimental and the of the the to cross reporter ion interference, in quantification accuracy by proteomic using TMT become of the to low missing values and precision when a single multiplexed batch is when large that the of multiple TMT batches are J. J. of tandem mass tags (TMT) and high specific proteome in 2017; Full Text Full Text PDF PubMed Scopus Google Scholar, J. normalization method for by 15: Full Text Full Text PDF PubMed Scopus Google the more we have used the of data integrated from 10-plex TMT batches to missing false coisolation interference, reporter ion interference and experimental design within large-scale proteomics We have on a data set derived from the of human cell derived from both male and female A. A. S. S. A. A. R. F. Mann A. A. R. S. M. A. R. A. R. O. in human 2017; PubMed Scopus Google Scholar). The data that a single batch of TMT the missing values in proteomics with data dependent acquisition (DDA), both at the protein and peptide this as data from or more TMT batches are integrated. multiple batches are the missing values effect is at the peptide integrating data from batches the missing values to from to the inflation rate at the protein level is the of the batch missing protein values from to effect the accuracy of results derived from large-scale that data from multiple TMT batches. be to TMT as this to more peptide S. S. R. J. J. of for tandem mass quantitative proteomics across 2018; Google is this more across batches. TMT quantitation the effect of the coisolation interference. single TMT batches results within the multiplexed we that this is not across multiple batches. study we the data using the proteomic Cox J. Mann M. for protein and without Full Text Full Text PDF PubMed Scopus Google and for every protein we the of both for the of the reference and within of 10-plex TMT batches. The of the was data from within the same 10-plex TMT also the of batch effects, in via a common within TMT batch for data normalization to minimize the batch effects, as of large of human using Biol. 2017; PubMed Scopus Google Scholar, M. quantitative of the human proteome in and 2018; PubMed Scopus Google Scholar). a of data are to address batch We that by at within TMT batch the batch be The in a that is for proteins within the experiment and in a that is across all the TMT batches. study also highlighted the of false reporter ion interference and coisolation interference. The data set we an set to these as iPSC lines derived from both male and female by a set of peptides mapped to the Y these a set of internal to the of false The data that for a 10-plex TMT batch with male channels the female channels quantified of all the Y chromosome-specific peptides that were detected in that are false positives consistently detected in the female channels within the multiplexed are limitations when within the same TMT we to coisolation interference and we that the not with the the to for ion interference in PubMed Scopus Google Scholar). new have to be coisolation Winter S. F. C. Cox J. Mann M. F. enables multiplexed and proteome Methods. 2018; 15: PubMed Scopus Google is to a We used the Y chromosome peptides to study the reporter ion interference across 10-plex TMT batches with of male and female derived cell as well as The data highlighted the of primary into a female and secondary into a female reporter ion interference. channels by both primary and secondary reporter ion interference a in high peptides of with channels not to reporter ion interference. was to be caused by as the data also that the effect with a with channels with reporter ion interference. the of reporter ion interference, we have used these data to propose experimental set for to specific channels that or the effect of primary and secondary reporter ion interference Nonetheless, we that within a TMT for and false positives within the as by the Y chromosome-specific peptides detected within all female cell lines. large-scale is also to have in to evaluate and a within the of the TMT batches in the The was not detected the were results within that We that to large-scale TMT a should be set in the of the TMT is a for and to and quantitation have a for proteomic we have an of the of quantitative data from large-scale proteomics and we of the limitations should be when these We the for experimental design and data for proteomics The used for this are at R. A. F. M. A. S. R. proteomics data and PubMed Scopus Google via the A. J. F. R. of the and PubMed Scopus Google with the with the J. Mann M. enables high peptide identification mass and protein PubMed Scopus Google We to the of the for the and of the for the with
