Graphlet Kernels for Prediction of Functional Residues in Protein Structures

We introduce a novel graph-based kernel method for annotating functional residues in protein structures. A structure is first modeled as a protein contact graph, where nodes correspond to residues and edges connect spatially neighboring residues. Each vertex in the graph is then represented as a vector of counts of labeled non-isomorphic subgraphs (graphlets), centered on the vertex of interest. A similarity measure between two vertices is expressed as the inner product of their respective count vectors and is used in a supervised learning framework to classify protein residues. We evaluated our method on two function prediction problems: identification of catalytic residues in proteins, which is a well-studied problem suitable for benchmarking, and a much less explored problem of predicting phosphorylation sites in protein structures. The performance of the graphlet kernel approach was then compared against two alternative methods, a sequence-based predictor and our implementation of the FEATURE framework. On both tasks, the graphlet kernel performed favorably; however, the margin of difference was considerably higher on the problem of phosphorylation site prediction. While there is data that phosphorylation sites are preferentially positioned in intrinsically disordered regions, we provide evidence that for the sites that are located in structured regions, neither the surface accessibility alone nor the averaged measures calculated from the residue microenvironments utilized by FEATURE were sufficient to achieve high accuracy. The key benefit of the graphlet representation is its ability to capture neighborhood similarities in protein structures via enumerating the patterns of local connectivity in the corresponding labeled graphs.

Derivation of 3Dcoordinate templates fo…Derivation of 3D coordinate templates for searching structural databases: Application to ser‐His‐Asp catalytic triads in the serine proteinases and lipasesTess: A geometrichashing algorithm for…Tess: A geometric hashing algorithm for deriving 3D coordinate templates for searching structural databases. Application to enzyme active sitesRecognition of spatialmotifs in protein…Recognition of spatial motifs in protein structures 1 1Edited by J. ThorntonFunctional Sites inProtein Families…Functional Sites in Protein Families Uncovered via an Objective and Automated Graph Theoretic ApproachAutomated prediction ofprotein function and…Automated prediction of protein function and detection of functional sites from structureProtein FunctionPrediction Using Local…Protein Function Prediction Using Local 3D TemplatesInference of ProteinFunction from Protein…Inference of Protein Function from Protein StructurePredicting proteinfunction from sequence…Predicting protein function from sequence and structural dataProFunc: a server forpredicting protein…ProFunc: a server for predicting protein function from 3D structureGraph kernels forchemical informaticsGraph kernels for chemical informaticsPredicting proteinfunction from sequence…Predicting protein function from sequence and structurePHOSIDA (phosphorylationsite database)…PHOSIDA (phosphorylation site database): management, structural and evolutionary investigation, and prediction of phosphositesStructure-based kernelsfor the prediction of…Structure-based kernels for the prediction of catalytic residues and their involvement in human inherited diseaseGUISE: Uniform Samplingof Graphlets for Large…GUISE: Uniform Sampling of Graphlets for Large Graph AnalysisGeneralized graphletkernels for…Generalized graphlet kernels for probabilistic inference in sparse graphsExploring the structureand function of tempora…Exploring the structure and function of temporal networks with dynamic graphletsThe Loss and Gain ofFunctional Amino Acid…The Loss and Gain of Functional Amino Acid Residues Is a Common Mechanism Causing Human Inherited DiseaseE-CLoG: Countingedge-centric local…E-CLoG: Counting edge-centric local graphletsUltra High-DimensionalNonlinear Feature…Ultra High-Dimensional Nonlinear Feature Selection for Big Biological DataClassification inbiological networks wit…Classification in biological networks with hypergraphlet kernelsA Learned Sketch forSubgraph CountingA Learned Sketch for Subgraph CountingNeural Subgraph Countingwith Wasserstein…Neural Subgraph Counting with Wasserstein EstimatorLearned sketch forsubgraph counting: a…Learned sketch for subgraph counting: a holistic approachCurrent and futuredirections in network…Current and future directions in network biologyGraphlet Kernels forPrediction of Functiona…Graphlet Kernels for Prediction of Functional Residues in Protein StructuresEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.