Dimensionality Reduction for Supervised Learning with Reproducing Kernel Hilbert Spaces

We propose a novel method of dimensionality reduction for supervised learning problems. Given a regression or classification problem in which we wish to predict a response variable $Y$ from an explanatory variable $X$, we treat the problem of dimensionality reduction as that of finding a low-dimensional ``effective subspace'' of $X$ which retains the statistical relationship between $X$ and $Y$. We show that this problem can be formulated in terms of conditional independence. To turn this formulation into an optimization problem we establish a general nonparametric characterization of conditional independence using covariance operators on a reproducing kernel Hilbert space. This characterization allows us to derive a contrast function for estimation of the effective subspace. Unlike many conventional methods for dimensionality reduction in supervised learning, the proposed method requires neither assumptions on the marginal distribution of $X$, nor a parametric model of the conditional distribution of $Y$. We present experiments that compare the performance of the method with conventional methods.

Joint measures andcross-covariance…Joint measures and cross-covariance operatorsGeneralized AdditiveModelsGeneralized Additive ModelsBayesian VariableSelection in Linear…Bayesian Variable Selection in Linear RegressionSliced InverseRegression for Dimensio…Sliced Inverse Regression for Dimension ReductionSliced InverseRegression for Dimensio…Sliced Inverse Regression for Dimension Reduction: CommentOn Principal HessianDirections for Data…On Principal Hessian Directions for Data Visualization and Dimension Reduction: Another Application of Stein's LemmaVariable Selection ViaGibbs SamplingVariable Selection Via Gibbs SamplingNeural Networks forPattern RecognitionNeural Networks for Pattern RecognitionRegression Shrinkage andSelection Via the LassoRegression Shrinkage and Selection Via the LassoNonlinear ComponentAnalysis as a Kernel…Nonlinear Component Analysis as a Kernel Eigenvalue ProblemBayesian modelaveraging: a tutorial…Bayesian model averaging: a tutorial (with comments by M. Clyde, David Draper and E. I. George, and a rejoinder by the authorsTheory & Methods:Special Invited Paper…Theory & Methods: Special Invited Paper: Dimension Reduction and Visualization in Discriminant Analysis (with discussion)Regression on manifoldsusing kernel dimension…Regression on manifolds using kernel dimension reductionGene selection via theBAHSIC family of…Gene selection via the BAHSIC family of algorithmsKernel dimensionreduction in regressionKernel dimension reduction in regressionTwo Manifold Problemswith Applications to…Two Manifold Problems with Applications to Nonlinear System IdentificationFeature-aware LabelSpace Dimension…Feature-aware Label Space Dimension Reduction for Multi-label ClassificationGENERALIZED DOUBLEPARETO SHRINKAGE.GENERALIZED DOUBLE PARETO SHRINKAGE.A GENERAL THEORY FORNONLINEAR SUFFICIENT…A GENERAL THEORY FOR NONLINEAR SUFFICIENT DIMENSION REDUCTION: FORMULATION AND ESTIMATIONSufficient Reductions inRegressions With…Sufficient Reductions in Regressions With Exponential Family Inverse PredictorsKernel Mean ShrinkageEstimatorsKernel Mean Shrinkage EstimatorsEstimation of COVID-19spread curves…Estimation of COVID-19 spread curves integrating global data and borrowing informationSpike-and-Slab MeetsLASSO: A Review of the…Spike-and-Slab Meets LASSO: A Review of the Spike-and-Slab LASSOBayesian approaches tovariable selection: a…Bayesian approaches to variable selection: a comparative study from practical perspectivesDimensionality Reductionfor Supervised Learning…Dimensionality Reduction for Supervised Learning with Reproducing Kernel Hilbert SpacesEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.