Empowering large chemical knowledge bases for exposomics: PubChemLite meets MetFrag
Keynote Presentation for the MS in Data Science Session at the IMSC2022, Maastricht Empowering Large Chemical Knowledge Bases for Exposomics: PubChemLite meets MetFrag Emma L. Schymanski1*, Todor Kondic1, Steffen Neumann2, Paul Thiessen3, Jian Zhang3, Evan E. Bolton3 1Luxembourg Centre for Systems Biomedicine (LCSB), University of Luxembourg, 6 avenue du Swing, 4367 Belvaux, Luxembourg. 2Leibniz Institute of Plant Biochemistry (IPB Halle), Bioinformatics and Scientific Data, 06120 Halle, Germany and German Centre for Integrative Biodiversity Research (iDiv), Halle-Jena-Leipzig, Deutscher Platz 5e, 04103 Leipzig, Germany. 3National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, MD 20894, USA. Introduction Exposomics researchers need to identify relevant chemicals covering the entirety of potential exposures over entire lifetimes. With over 100 million chemicals in the largest chemical databases, coupled with broadly acknowledged knowledge gaps, researchers are faced with too much—yet not enough—information at the same time. Improvements in analytical technologies and computational mass spectrometry workflows coupled with the rapid growth in databases and increasing demand for high throughput “big data” services from the research community present significant challenges for both data hosts and workflow developers. The “PubChemLite for Exposomics” collection reduces candidate search spaces in non-target small molecule identification workflows while increasing content usability. This allows users to profit from both increasing size and information content of large compound databases, along with increased efficiency. Methods PubChemLite for Exposomics is a dynamic collection of ~380,000 chemicals that is built weekly from several categories of annotation content in the PubChem database that are highly relevant for exposomics analysis. These ten categories are: Agrochemical Information; Associated Disorders and Diseases; Biomolecular Interactions and Pathways; Drug and Medication Information; Food Additives and Ingredients; Identification; Pharmacology and Biochemistry; Safety and Hazards; Toxicity; Use and Manufacturing. Benchmarking datasets are used to show how experimental knowledge and existing datasets can help detect and fill gaps in compound databases to progressively improve large resources such as PubChem, and topic-specific subsets such as PubChemLite. Preliminary data (results) Interdisciplinary efforts and data sharing can facilitate research in exposomics and beyond. This effort demonstrates this using the examples of PubChem (https://pubchem.ncbi.nlm.nih.gov/), the NORMAN Network Suspect List Exchange (https://www.norman-network.com/nds/SLE/) and the in silico fragmentation approach MetFrag (https://msbi.ipb-halle.de/MetFrag/). A subset of the PubChem database relevant for exposomics, PubChemLite for Exposomics, is presented as a database resource that can be (and has been) integrated into current workflows for high resolution mass spectrometry. The benchmarking analysis performed demonstrated dramatic performance improvements over MetFrag coupled with PubChem, both in terms of candidate ranking (improving ranking to 81 % candidates correct in first place) and runtime. PubChemLite for Exposomics is a living collection, updating as annotation content in PubChem is updated (exemplified using the NORMAN Suspect List Exchange), and exported to allow direct integration into existing workflows such as MetFrag. The dataset is currently built and checked weekly and updated monthly on Zenodo (DOI: 10.5281/zenodo.5995885). It is integrated into the MetFrag Web interface (https://msbi.ipb-halle.de/MetFrag/) and available for download for command line use. It is also integrated into the open source mass spectrometry workflow patRoon (https://rickhelmus.github.io/patRoon/). The source code and files necessary to create and adjust this collection are jointly hosted between the research parties. This effort shows that enhancing the FAIRness (Findability, Accessibility, Interoperability and Reusability) of open resources can mutually enhance several resources for whole community benefit. Mass spectrometry related innovations PubChemLite for Exposomics is a dynamic collection of ~380,000 chemicals formed from several annotation categories in PubChem that is designed to empower exposomics analysis using computational high resolution mass spectrometry.
