Exploring Large Feature Spaces with Hierarchical Multiple Kernel Learning

For supervised and unsupervised learning, positive definite kernels allow to use large and potentially infinite dimensional feature spaces with a computational cost that only depends on the number of observations. This is usually done through the penalization of predictor functions by Euclidean or Hilbertian norms. In this paper, we explore penalizing by sparsity-inducing norms such as the l1-norm or the block l1-norm. We assume that the kernel decomposes into a large sum of individual basis kernels which can be embedded in a directed acyclic graph; we show that it is then possible to perform kernel selection through a hierarchical multiple kernel learning framework, in polynomial time in the number of selected kernels. This framework is naturally applied to non linear variable selection; our extensive simulations on synthetic datasets and datasets from the UCI repository show that efficiently exploring the large feature space through sparsity-inducing norms leads to state-of-the-art predictive performance.

Computing regularizationpaths for learning…Computing regularization paths for learning multiple kernelsMultiple kernellearning, conic duality…Multiple kernel learning, conic duality, and the SMO algorithmConvex OptimizationConvex OptimizationKernel Methods forPattern AnalysisKernel Methods for Pattern AnalysisLearning the KernelFunction via…Learning the Kernel Function via RegularizationLarge Scale MultipleKernel LearningLarge Scale Multiple Kernel LearningThe Adaptive Lasso andIts Oracle PropertiesThe Adaptive Lasso and Its Oracle PropertiesOn Model SelectionConsistency of LassoOn Model Selection Consistency of LassoMore efficiency inmultiple kernel learningMore efficiency in multiple kernel learningGrouped and HierarchicalModel Selection through…Grouped and Hierarchical Model Selection through Composite Absolute PenaltiesConsistency of the GroupLasso and Multiple…Consistency of the Group Lasso and Multiple Kernel LearningComposite kernellearningComposite kernel learningLearning Non-LinearCombinations of KernelsLearning Non-Linear Combinations of KernelsMore generality inefficient multiple…More generality in efficient multiple kernel learningGeneralization Boundsfor Learning KernelsGeneralization Bounds for Learning KernelsTwo-Stage LearningKernel AlgorithmsTwo-Stage Learning Kernel AlgorithmsNon-SparseRegularization and…Non-Sparse Regularization and Efficient Training with Multiple KernelsStructuredsparsity-inducing norms…Structured sparsity-inducing norms through submodular functionslp-Norm Multiple KernelLearninglp-Norm Multiple Kernel LearningGroup Lasso withOverlaps: the Latent…Group Lasso with Overlaps: the Latent Group Lasso approachStructured sparsitythrough convex…Structured sparsity through convex optimizationMultiple Kernel Learningfor Visual Object…Multiple Kernel Learning for Visual Object Recognition: A ReviewNonlinear Deep KernelLearning for Image…Nonlinear Deep Kernel Learning for Image AnnotationMultiple Kernel Learningfor Remote Sensing Imag…Multiple Kernel Learning for Remote Sensing Image ClassificationExploring Large FeatureSpaces with Hierarchica…Exploring Large Feature Spaces with Hierarchical Multiple Kernel LearningEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.