xiNET: Cross-link Network Maps With Residue Resolution
xiNET is a visualization tool for exploring cross-linking/mass spectrometry results. The interactive maps of the cross-link network that it generates are a type of node-link diagram. In these maps xiNET displays: (1) residue resolution positional information including linkage sites and linked peptides; (2) all types of cross-linking reaction product; (3) ambiguous results; and, (4) additional sequence information such as domains. xiNET runs in a browser and exports vector graphics which can be edited in common drawing packages to create publication quality figures. Availability: xiNET is open source, released under the Apache version 2 license. Results can be viewed by uploading data to http://crosslinkviewer.org/ or by downloading the software from http://github.com/colin-combe/crosslink-viewer and running it locally. xiNET is a visualization tool for exploring cross-linking/mass spectrometry results. The interactive maps of the cross-link network that it generates are a type of node-link diagram. In these maps xiNET displays: (1) residue resolution positional information including linkage sites and linked peptides; (2) all types of cross-linking reaction product; (3) ambiguous results; and, (4) additional sequence information such as domains. xiNET runs in a browser and exports vector graphics which can be edited in common drawing packages to create publication quality figures. Availability: xiNET is open source, released under the Apache version 2 license. Results can be viewed by uploading data to http://crosslinkviewer.org/ or by downloading the software from http://github.com/colin-combe/crosslink-viewer and running it locally. Cross-linking/mass spectrometry (CLMS) 1The abbreviations used are:CLMSCross-linking/mass spectrometryCSVComma Separated ValuesPDBProtein Data BankPSI-MIProteomics Standards Initiative – Molecular InteractionsRINResidue Interaction NetworkSVGScalable Vector Graphic. 1The abbreviations used are:CLMSCross-linking/mass spectrometryCSVComma Separated ValuesPDBProtein Data BankPSI-MIProteomics Standards Initiative – Molecular InteractionsRINResidue Interaction NetworkSVGScalable Vector Graphic. has revealed protein–protein interactions in large multiprotein complexes (1Chen Z.A. Jawhari A. Fischer L. Buchen C. Tahir S. Kamenski T. Rasmussen M. Lariviere L. Bukowski-Wills J.C. Nilges M. Cramer P. Rappsilber J. Architecture of the RNA polymerase II–TFIIF complex revealed by cross-linking and mass spectrometry.EMBO J. 2010; 29: 717-726Crossref PubMed Scopus (324) Google Scholar), small networks (2Herzog F. Kahraman A. Boehringer D. Mak R. Bracher A. Walzthoeni T. Leitner A. Beck M. Hartl F.U. Ban N. Malmström L. Aebersold R. Structural probing of a protein phosphatase 2a network by chemical cross-linking and mass spectrometry.Science. 2012; 337: 1348-1352Crossref PubMed Scopus (312) Google Scholar), and complex mixtures (3Zheng C. Yang L. Hoopmann M.R. Eng J.K. Tang X. Weisbrod C.R. Bruce J.E. Cross-linking measurements of in vivo protein complex topologies.Mol. Cell. Proteomics. 2011; 10M110.006841–M110.006841Abstract Full Text Full Text PDF Scopus (80) Google Scholar). When cross-links are observed they define a pair of residues that are close in space and in this way reveal not just the identity of interacting proteins but also pinpoint interacting regions or domains and their orientation. A common feature of CLMS studies is identifying many individual cross-links that cumulatively support the existence of domain-level features. Making this connection between the network of residue distance constraints and domain-level features is difficult if looking at a large table of CLMS data. A visual analysis that highlights clusters of cross-links within the protein sequence helps greatly. As we shall see (“Node Layouts”), not all visualizations of the network will allow this. Cross-linking/mass spectrometry Comma Separated Values Protein Data Bank Proteomics Standards Initiative – Molecular Interactions Residue Interaction Network Scalable Vector Graphic. Cross-linking/mass spectrometry Comma Separated Values Protein Data Bank Proteomics Standards Initiative – Molecular Interactions Residue Interaction Network Scalable Vector Graphic. There are two types of network visualization: node-link diagrams and adjacency matrices. Each has its own advantages and disadvantages. Node-link diagrams preserve the local detail of the network but do not scale well. Adjacency matrices avoid the edge crossings that make large node-link diagrams unreadable, but make it difficult to understand the relationships between nodes that are not directly connected (4Gehlenborg N. Wong B. Points of view: Networks.Nat. Methods. 2012; 9: 115Crossref PubMed Scopus (10) Google Scholar). Seebacher et al. (5Seebacher J. Mallick P. Zhang N. Eddes J.S. Aebersold R. Gelb M.H. Protein cross-linking analysis using mass spectrometry, isotope-coded cross-linkers, and integrated computational data processing.J. Proteome Res. 2006; 5: 2270-2282Crossref PubMed Scopus (103) Google Scholar) show CLMS data visualized as an adjacency matrix. Here, we focus on visualizing CLMS data as a node-link diagram and present xiNET, a network visualization tool designed specifically for use with CLMS data. In 2010, Gehlenborg et al. (6Gehlenborg N. O'Donoghue S.I. Baliga N.S. Goesmann A. Hibbs M.A. Kitano H. Kohlbacher O. Neuweger H. Schneider R. Tenenbaum D. Gavin A.C. Visualization of omics data for systems biology.Nat. Methods. 2010; 7: S56-S68Crossref PubMed Scopus (462) Google Scholar) provided a review of the then state-of-the-art visualization tools for “omics” data in systems biology. They begin by noting that all these tools are dominated by the same primary visual metaphor: graphs displayed as node-link diagrams. Their review is, essentially, a review of the use of node-link diagrams in biology. They continue by identifying two broad, partly overlapping, categories of visualization tools–pathway tools and network tools. Pathway tools display a graph representing changes in state over time; network tools display graphs that do not (necessarily) include state change information. CLMS can provide data on conformational changes within proteins, hence state change is one aspect of this data. However, we will narrow our focus here by concentrating on network tools that do not set out to represent state change. Gehlenborg et al. present a second categorization of network tools corresponding to the three main types of high throughput experiment. The categories are: tools for investigating protein-protein interactions, tools for investigating gene expression profiles, and tools for investigating metabolic profiles. They list 27 software packages specifically intended for investigating protein-protein interaction networks–a list that has continued to grow since 2010. Of these 27, they recommend two–CytoScape (7Shannon P. Markiel A. Ozier O. Baliga N.S. Wang J.T. Ramage D. Amin N. Schwikowski B. Ideker T. Cytoscape: A software environment for integrated models of biomolecular interaction networks.Genome Res. 2003; 13: 2498-2504Crossref PubMed Scopus (25650) Google Scholar) and Cerebral (8Barsky A. Gardy J.L. Hancock R.E. Munzner T. Cerebral: a Cytoscape plugin for layout of and interaction with biological networks using subcellular localization annotation.Bioinformatics. 2007; 23: 1040-1042Crossref PubMed Scopus (144) Google Scholar). CytoScape is perhaps the most popular network visualization tool in biology. The CytoScape software provides a plugin architecture that allows extension and customization. For example, a new node layout algorithm could be added. Cerebral is a CytoScape plugin that uses additional annotation information, particularly subcellular location, to guide the layout of the nodes. Its aim is to produce interaction network diagrams that more closely resemble “traditional” signaling pathway/system diagrams, with extracellular proteins and membrane receptors at the top of the page, moving down through adapter proteins in the cytoplasm, with nuclear proteins and pathway-regulated genes at the bottom. In biology, node-link diagrams typically use the nodes to represent entire molecules (4Gehlenborg N. Wong B. Points of view: Networks.Nat. Methods. 2012; 9: 115Crossref PubMed Scopus (10) Google Scholar). Much of the discussion of protein–protein interaction tools in Gehlenborg et al.'s 2010 review focuses on arranging the nodes that represent molecules according to the higher order structures (complexes and groups of complexes) that they make up. Related to this are approaches that display a hierarchical graph and in which nodes representing individual molecules can be collapsed into a single meta-node representing a higher order grouping. An example of such software is Visant (9Hu Z. Mellor J. Wu J. Kanehisa M. Stuart J.M. DeLisi C. Towards zoomable multidimensional maps of the cell.Nat. Biotechnol. 2007; 25: 547-554Crossref PubMed Scopus (73) Google Scholar). When discussing future directions for the network visualization of omics data, Gehlenborg et al. highlight improved navigation methods for large networks, a trend toward web-based tools, and the need for standardization of data formats. PSI-MI (10Hermjakob H. Montecchi-Palazzi L. Bader G. Wojcik J. Salwinski L. Ceol A. Moore S. Orchard S. Sarkans U. Von Mering C. Roechert B. Poux S. Jung E. Mersch H. Kersey P. Lappe M. Li Y. Zeng R. Rana D. Nikolski M. Husi H. Brun C. Shanker K. Grant S.G. Sander C. Bork P. Zhu W. Pandey A. Brazma A. Jacq B. Vidal M. Sherman D. Legrain P. Cesareni G. Xenarios I. Eisenberg D. Steipe B. Hogue C. Apweiler R. The HUPO PSI's molecular interaction format–a community standard for the of protein interaction Biotechnol. PubMed Scopus Google Scholar) is with the standardization of interaction data. However, in their review is a discussion of the nodes down into An to representing as nodes is to use nodes to represent and are tools that do this. is in tools for residue interaction networks from protein data J. Z. G. H. The protein data Res. PubMed Scopus Google Scholar) Y. M. analysis and interactive visualization of biological networks and protein 2012; 7: PubMed Scopus Google Scholar) is an example of such a is used to conformational changes example, as a of or looking at residue relationships of information through a uses the to guide the layout of the nodes in the network diagram. structures or models are in the of CLMS data, for most proteins we do not such data and in of sequence space is types of data such as is as a CytoScape to is M. F. T. I. interacting information and in protein 2011; PubMed Scopus Google Scholar), it also generates a from a and this a CytoScape to use the to guide the node layout in A tool we of is L. of as PubMed Scopus Google Scholar). a CytoScape but a the interacting residues an representing the protein However, is by the visualization of interactions between two of the is, it can display two tools the most in common with our for visualizing CLMS data, as it is a network of interacting residues that CLMS data CLMS data into tools, these tools do not our for visualizing such data. understand it is to at the of node layout in more that a biological use for CLMS data is drawing domain-level features from the many individual The residue distance constraints from a CLMS an graph and are many network visualization tools that could be used to display such a However, are for a visualization that allows clusters of linked residues within the protein sequence to be the of domain-level features. is in which node-link of the same CLMS data. For we use the cross-links from this to a protein of the same The data is from et al. (1Chen Z.A. Jawhari A. Fischer L. Buchen C. Tahir S. Kamenski T. Rasmussen M. Lariviere L. Bukowski-Wills J.C. Nilges M. Cramer P. Rappsilber J. Architecture of the RNA polymerase II–TFIIF complex revealed by cross-linking and mass spectrometry.EMBO J. 2010; 29: 717-726Crossref PubMed Scopus (324) Google Scholar), the of the RNA polymerase to their is CLMS for the of domains between and is a et Its to this for the of the domains. is a type of node-link diagram which the protein as and cross-links as at these However, are many of arranging the of which will an to the CLMS data for in the way most in biology, with nodes representing it to the of the domains within the protein sequence the information is However, are cross-linking studies that do their in this way (2Herzog F. Kahraman A. Boehringer D. Mak R. Bracher A. Walzthoeni T. Leitner A. Beck M. Hartl F.U. Ban N. Malmström L. Aebersold R. Structural probing of a protein phosphatase 2a network by chemical cross-linking and mass spectrometry.Science. 2012; 337: 1348-1352Crossref PubMed Scopus (312) Google K. F. S. Walzthoeni T. E. P. Beck F. Aebersold R. A. W. Molecular architecture of the by an 2012; PubMed Scopus Google Scholar). For visualizing the data at this of is A with CLMS data is the of two of residues and linked could be the nodes in a An to representing as is to use nodes to represent as is in and our example CLMS data in this Each linked residue is a but a used to guide the layout of the nodes not the of the domains in the protein of the this not is that the residue networks that from cross-link data are connected from data. of the of the nodes in a space M. I. M.A. to visualizing 2011; Google Scholar). In to this is with node-link diagrams and the of the nodes the data is (4Gehlenborg N. Wong B. Points of view: Networks.Nat. Methods. 2012; 9: 115Crossref PubMed Scopus (10) Google Scholar). The of the nodes in can be such that it will a to the of the domains. do the used to the nodes the linked residues they represent within the of the protein see and order the of the linked residue nodes a but is the of the within the protein When the their within the protein a M. I. M.A. to visualizing 2011; Google Scholar) of the data. the of using the of the linked residues within the protein sequence arranging the nodes. it in the connection between the CLMS and the of the domains within the In a the are categories of in these categories represent the three proteins in the data and the for the is on the residue within the are more three categories is, more three in our more three the linked residues representing the protein as is to interactions between two is the visualization we for CLMS data. the from et could be of as a as but with the and Here, we present tool for interactive of the of node-link diagram in layout has used in CLMS for (1Chen Z.A. Jawhari A. Fischer L. Buchen C. Tahir S. Kamenski T. Rasmussen M. Lariviere L. Bukowski-Wills J.C. Nilges M. Cramer P. Rappsilber J. Architecture of the RNA polymerase II–TFIIF complex revealed by cross-linking and mass spectrometry.EMBO J. 2010; 29: 717-726Crossref PubMed Scopus (324) Google S. F. Aebersold R. Cramer P. analysis RNA polymerase architecture and of Res. 2012; PubMed Scopus Google M.A. B. A. J. C. A. U. Rappsilber J. Structural for by the 5: PubMed Scopus Google A. A. L. T. W. A. J.S. Beck M. Structural of the Full Text Full Text PDF PubMed Scopus Google Scholar). is not the to visualizing this data that could show the many cross-links support the of domain-level features. There are of which for example, be as CytoScape However, the is not a for linked residue and allows the of the diagrams use this to closely connected the linked residues within the of the protein sequence xiNET an biological with CLMS data. However, xiNET has features this that will xiNET ambiguous the linkage can for example, an to more one protein in the ambiguous information can be or if not xiNET also and all cross-linking In to cross-linking and linked three types information. The of which of the protein An of in an they be can that the linked provide residue distance the of these from that of and they be in the An linked is to from an in which from the same protein could be or However, is a of in which the in the protein sequence and these are to be xiNET also these that could be from for the of the in CLMS is that it allows the to be in the of sequence information such as domains. xiNET such they can be and into the diagram. information with that the node-link diagrams in which nodes represent molecules are also a common network of CLMS data, and that this of be for xiNET allows these to be the data can be collapsed to protein node the or to show the linked residues a allows the to which of the network are in more or these features are in the Results xiNET the of CLMS data. this by it to the interactive and vector that can be used to create publication quality figures. can xiNET is in and a vector within a The interactive maps in can be The can be edited in common vector drawing such as or used M. J. Visualization 2011; PubMed Scopus Google Scholar) as a for common visualization such as The xiNET has an and the types make the a Residue of all the between a pair of Protein of all the residue between two and 2 provides an of the xiNET The data for a cross-link are: cross-link protein sequence and, annotation data. data can be if on at the Protein in Res. PubMed Scopus Google Scholar) are used as the protein in this protein will be using a provided by are used then annotation data will also be from this and from the J. K. R. C. of to using a of models that represent all proteins of PubMed Scopus Google Scholar) A. L. The annotation PubMed Scopus Google Scholar) data has into xiNET the CLMS network can be and vector graphics can be for use in figures. Data can be into xiNET by uploading it to our or by downloading the software and running it locally. of these are in the studies in the Results When used xiNET not of the data the A xiNET into its can of all data used and or to make data use xiNET as a data are to and of the are are to a their data. an interactive is then as as that The of xiNET for CLMS data is a will to the xiNET as is at the of the data typically in the information that cross-linking for see (1Chen Z.A. Jawhari A. Fischer L. Buchen C. Tahir S. Kamenski T. Rasmussen M. Lariviere L. Bukowski-Wills J.C. Nilges M. Cramer P. Rappsilber J. Architecture of the RNA polymerase II–TFIIF complex revealed by cross-linking and mass spectrometry.EMBO J. 2010; 29: 717-726Crossref PubMed Scopus (324) Google F. Kahraman A. Boehringer D. Mak R. Bracher A. Walzthoeni T. Leitner A. Beck M. Hartl F.U. Ban N. Malmström L. Aebersold R. Structural probing of a protein phosphatase 2a network by chemical cross-linking and mass spectrometry.Science. 2012; 337: 1348-1352Crossref PubMed Scopus (312) Google A. A. L. T. W. A. J.S. Beck M. Structural of the Full Text Full Text PDF PubMed Scopus Google Scholar). is also to the for C. Weisbrod C.R. Eng J.K. Wu X. Bruce J.E. and software tools for and visualizing protein interaction Proteome Res. PubMed Scopus Google Scholar) and the of O. Seebacher J. Walzthoeni T. Beck M. A. M. Aebersold R. of from large sequence Methods. 5: PubMed Scopus Google Scholar). are the standard for CLMS data and the can be However, the and for all these and tools be PSI-MI (10Hermjakob H. Montecchi-Palazzi L. Bader G. Wojcik J. Salwinski L. Ceol A. Moore S. Orchard S. Sarkans U. Von Mering C. Roechert B. Poux S. Jung E. Mersch H. Kersey P. Lappe M. Li Y. Zeng R. Rana D. Nikolski M. Husi H. Brun C. Shanker K. Grant S.G. Sander C. Bork P. Zhu W. Pandey A. Brazma A. Jacq B. Vidal M. Sherman D. Legrain P. Cesareni G. Xenarios I. Eisenberg D. Steipe B. Hogue C. Apweiler R. The HUPO PSI's molecular interaction format–a community standard for the of protein interaction Biotechnol. PubMed Scopus Google Scholar) and we are toward using this xiNET is designed to allow its use as a that can into their by it directly to their own CLMS In data can be into xiNET by through and to and xiNET we provide a tool that CLMS in an interactive tool is an open at and can be used and by The tool can be used directly through a at or used the data. The for the interactive maps the are in on xiNET linkage the proteins between and on on and on and on through protein and on that at of a protein is protein all its on between two on between all on of on which are on protein or detail in in protein or on over protein of linkage of over protein-protein of linked over residue and linked in table in a new In this we provide for a xiNET by uploading data to the Data to to all data within the by using the tool show these maps and then to all the features of the for xiNET, which is used in the a such as that in in a drawing tool is a xiNET this In we to an interactive version of from the data the publication and it with the domains et al. (1Chen Z.A. Jawhari A. Fischer L. Buchen C. Tahir S. Kamenski T. Rasmussen M. Lariviere L. Bukowski-Wills J.C. Nilges M. Cramer P. Rappsilber J. Architecture of the RNA polymerase II–TFIIF complex revealed by cross-linking and mass spectrometry.EMBO J. 2010; 29: 717-726Crossref PubMed Scopus (324) Google Scholar) the linkage the sequence information, in an the from the we a in the way the in or as these are not in the the as a the to at in a to the protein for the uses as is need to provide sequence data. An to be to provide sequence data in a the used as The annotation data is from the and in the at the and annotation are then The is in An interactive version is at The in feature of xiNET identifying There is a in the data residue to residue if could be and could from a xiNET this and highlights et al. A. A. L. T. W. A. J.S. Beck M. Structural of the Full Text Full Text PDF PubMed Scopus Google Scholar) the nuclear complex with They provided with O. Seebacher J. Walzthoeni T. Beck M. A. M. Aebersold R. of from large sequence Methods. 5: PubMed Scopus Google Scholar) this complex for use as example data and this data all three cross-linking reaction the to the and within as this information and it as the However, for the of xiNET directly as it is to its own the types are in they can also be from the of the interactive the cross-link data the xiNET from the in the in of from to in a to the be to allow it to from the local this not if the are by a running Network is as and are but the cross-link data not the local The are in in which we see xiNET and all three The version of the is at are in order to all the features of xiNET in the of ambiguous data displayed in are used to the ambiguous and highlights on show the In the for the sequence of the protein is When xiNET has information on the its highlights the and the linked on in we can see the that is the of at a aspect of of sequence features such a domains. at When proteins are as the domains are as regions this When the nodes are collapsed into a annotation information is on these nodes as the and of which to the and residues of the the are on the of their proteins can be by in their A feature of is its to nodes between a and a to protein allows the to which of the network are and which more detail aspect of xiNET is that it provides a hierarchical is to Visant (9Hu Z. Mellor J. Wu J. Kanehisa M. Stuart J.M. DeLisi C. Towards zoomable multidimensional maps of the cell.Nat. Biotechnol. 2007; 25: 547-554Crossref PubMed Scopus (73) Google Scholar), Visant not use an representing the sequence for nodes. that our will the of CLMS as an by and the way and the information provided by this type of experiment. the of interactive to and as for to the of a common visual for CLMS results. The open of xiNET the that its use standard the a of which from the biological we to xiNET not for all our visualizing of CLMS data. types of are in CLMS for cross-links visualized on a and a of to three types of in a for the visualization and analysis of CLMS data. xiNET is a one of these of for the of visualizing that for example, a web-based for Cell. Proteomics. 13: Full Text Full Text PDF PubMed Scopus Google Scholar) and at C. Weisbrod C.R. Eng J.K. Wu X. Bruce J.E. and software tools for and visualizing protein interaction Proteome Res. PubMed Scopus Google Scholar) a intended for use with CLMS data. xiNET provides a way of CLMS data and we it will a in an integrated that all of CLMS in an The maps that xiNET could be to include types of such as RNA or small The of the of between and could be to the visualization of data in which interactions are between regions of these the of types of and the of types of data, are of xiNET into a for interactions, see Bukowski-Wills for discussion the of the for on the and all of our for and of
