From Transformation-Based Dimensionality Reduction to Feature Selection

Many learning applications are characterized by high dimensions. Usually not all of these dimensions are relevant and some are redundant. There are two main approaches to reduce dimensionality: feature selection and feature transformation. When one wishes to keep the original meaning of the features, feature selection is desired. Feature selection and transformation are typically presented separately. In this paper, we introduce a general approach for converting transformation-based methods to feature selection methods through l1/l∞ regularization. Instead of solving feature selection as a discrete optimization, we relax and formulate the problem as a continuous optimization problem. An additional advantage of our formulation is that our optimization criterion optimizes for feature relevance and redundancy removal automatically. Here, we illustrate how our approach can be utilized to convert linear discriminant analysis (LDA) and the dimensionality reduction version of the Hilbert-Schmidt Independence Criterion (HSIC) to two new feature selection algorithms. Experiments show that our new feature selection methods out-perform related state-of-the-art feature selection approaches.

Regression Shrinkage andSelection Via the LassoRegression Shrinkage and Selection Via the LassoWrappers for FeatureSubset SelectionWrappers for Feature Subset SelectionUCI Repository ofmachine learning…UCI Repository of machine learning databasesStatistical LearningTheoryStatistical Learning Theory10.1162/15324430332275361610.1162/153244303322753616Gene Selection forCancer Classification…Gene Selection for Cancer Classification using Support Vector MachinesLaplacian Score forFeature SelectionLaplacian Score for Feature SelectionFeature Selection Basedon Mutual Information…Feature Selection Based on Mutual Information: Criteria of Max-Dependency, Max-Relevance, and Min-RedundancyModel Selection andEstimation in Regressio…Model Selection and Estimation in Regression with Grouped VariablesMeasuring StatisticalDependence with…Measuring Statistical Dependence with Hilbert-Schmidt NormsGeneralized spectralbounds for sparse LDAGeneralized spectral bounds for sparse LDASupervised FeatureSelection via Dependenc…Supervised Feature Selection via Dependence EstimationJoint Feature Selectionand Subspace LearningJoint Feature Selection and Subspace LearningLinear DiscriminantDimensionality ReductionLinear Discriminant Dimensionality ReductionFeature Selection viaL1-Penalized…Feature Selection via L1-Penalized Squared-Loss Mutual InformationRobust UnsupervisedFeature SelectionRobust Unsupervised Feature SelectionJoint Embedding Learningand Sparse Regression…Joint Embedding Learning and Sparse Regression: A Framework for Unsupervised Feature SelectionFeature Selection forClassification: A ReviewFeature Selection for Classification: A ReviewSparse discriminativefeature selectionSparse discriminative feature selectionEffective DiscriminativeFeature Selection With…Effective Discriminative Feature Selection With Nontrivial SolutionKernel Feature Selectionvia Conditional…Kernel Feature Selection via Conditional Covariance MinimizationAdaptive UnsupervisedFeature Selection With…Adaptive Unsupervised Feature Selection With Structure RegularizationLocal AdaptiveProjection Framework fo…Local Adaptive Projection Framework for Feature Selection of Labeled and Unlabeled DataSupervised FeatureSelection With a…Supervised Feature Selection With a Stratified Feature Weighting MethodFromTransformation-Based…From Transformation-Based Dimensionality Reduction to Feature SelectionEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.