Fast Learning Rate of Multiple Kernel Learning: Trade-Off between Sparsity and Smoothness

We investigate the learning rate of multiple kernel learning (MKL) with $\ell_{1}$ and elastic-net regularizations. The elastic-net regularization is a composition of an $\ell_{1}$-regularizer for inducing the sparsity and an $\ell_{2}$-regularizer for controlling the smoothness. We focus on a sparse setting where the total number of kernels is large, but the number of nonzero components of the ground truth is relatively small, and show sharper convergence rates than the learning rates have ever shown for both $\ell_{1}$ and elastic-net regularizations. Our analysis reveals some relations between the choice of a regularization function and the performance. If the ground truth is smooth, we show a faster convergence rate for the elastic-net regularization with less conditions than $\ell_{1}$-regularization; otherwise, a faster convergence rate for the $\ell_{1}$-regularization is shown.

Learning the KernelFunction via…Learning the Kernel Function via RegularizationLearning Bounds forSupport Vector Machines…Learning Bounds for Support Vector Machines with Learned KernelsA DC-programmingalgorithm for kernel…A DC-programming algorithm for kernel selectionLearning Non-LinearCombinations of KernelsLearning Non-Linear Combinations of KernelsL2 Regularization forLearning KernelsL2 Regularization for Learning KernelsHigh-dimensionaladditive modelingHigh-dimensional additive modelingMore generality inefficient multiple…More generality in efficient multiple kernel learningEfficient and AccurateLp-Norm Multiple Kernel…Efficient and Accurate Lp-Norm Multiple Kernel LearningSparsity-accuracytrade-off in MKLSparsity-accuracy trade-off in MKLA Unifying View ofMultiple Kernel LearningA Unifying View of Multiple Kernel LearningUnifying Framework forFast Learning Rate of…Unifying Framework for Fast Learning Rate of Non-Sparse Multiple Kernel LearningMinimax-Optimal RatesFor Sparse Additive…Minimax-Optimal Rates For Sparse Additive Models Over Kernel Classes Via Convex ProgrammingMinimax-Optimal RatesFor Sparse Additive…Minimax-Optimal Rates For Sparse Additive Models Over Kernel Classes Via Convex ProgrammingPAC-Bayesian Bound forGaussian Process…PAC-Bayesian Bound for Gaussian Process Regression and Multiple Kernel Additive ModelHigh-Dimensional FeatureSelection by…High-Dimensional Feature Selection by Feature-Wise Kernelized LassoMinimax Optimal Rates ofEstimation in High…Minimax Optimal Rates of Estimation in High Dimensional Additive Models: Universal Phase TransitionLearning rates for therisk of kernel-based…Learning rates for the risk of kernel-based quantile regression estimators in additive modelsMinimax optimal rates ofestimation in high…Minimax optimal rates of estimation in high dimensional additive modelsSparsity and ErrorAnalysis of Empirical…Sparsity and Error Analysis of Empirical Feature-Based Regularization SchemesUltra High-DimensionalNonlinear Feature…Ultra High-Dimensional Nonlinear Feature Selection for Big Biological DataA reproducing kernelHilbert space approach…A reproducing kernel Hilbert space approach to high dimensional partially varying coefficient modelImproved Learning Ratesof a Functional…Improved Learning Rates of a Functional Lasso-type SVM with Sparse Multi-Kernel RepresentationAccurate cancerphenotype prediction…Accurate cancer phenotype prediction with AKLIMATE, a stacked kernel learner integrating multimodal genomic data and pathway knowledgeGrouped VariableSelection with Discrete…Grouped Variable Selection with Discrete Optimization: Computational and Statistical PerspectivesFast Learning Rate ofMultiple Kernel…Fast Learning Rate of Multiple Kernel Learning: Trade-Off between Sparsity and SmoothnessEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.