Feature Selection for High-Dimensional Data: A Fast Correlation-Based Filter Solution

Feature selection, as a preprocessing step to machine learning, has been effective in reducing dimensionality, removing irrelevant data, increasing learning accuracy, and improving comprehensibility. However, the recent increase of dimensionality of data poses a severe challenge to many existing feature selection methods with respect to efficiency and effectiveness. In this work, we introduce a novel concept, predominant correlation, and propose a fast filter method which can identify relevant features as well as redundancy among relevant features without pairwise correlation analysis. The efficiency and effectiveness of our method is demonstrated through extensive comparisons with other methods using real-world data of high dimensionality. 1.

The Feature SelectionProblem: Traditional…The Feature Selection Problem: Traditional Methods and a New AlgorithmC4.5: Programs forMachine LearningC4.5: Programs for Machine LearningEstimating Attributes:Analysis and Extensions…Estimating Attributes: Analysis and Extensions of RELIEFSelection of RelevantFeatures in Machine…Selection of Relevant Features in Machine Learning.Wrappers for FeatureSubset SelectionWrappers for Feature Subset SelectionFeature Selection forClassificationFeature Selection for ClassificationFeature Selection forKnowledge Discovery and…Feature Selection for Knowledge Discovery and Data MiningCorrelation-basedFeature Selection for…Correlation-based Feature Selection for Machine LearningCorrelation-basedFeature Selection for…Correlation-based Feature Selection for Discrete and Numeric Class Machine LearningFilters, Wrappers and aBoosting-Based Hybrid…Filters, Wrappers and a Boosting-Based Hybrid for Feature SelectionUnsupervised FeatureSelection Using Feature…Unsupervised Feature Selection Using Feature SimilarityData Mining: PracticalMachine Learning Tools…Data Mining: Practical Machine Learning Tools and TechniquesBayesian Neural Networksfor Internet Traffic…Bayesian Neural Networks for Internet Traffic ClassificationSearching forinteracting features in…Searching for interacting features in subset selectionOnline feature selectionfor mining big dataOnline feature selection for mining big dataSelecting Feature Subsetvia Constraint…Selecting Feature Subset via Constraint Association RulesMeasuring stability offeature ranking…Measuring stability of feature ranking techniques: a noise-based approachRecent advances andemerging challenges of…Recent advances and emerging challenges of feature selection in the context of big dataDistributed featureselection: An…Distributed feature selection: An application to microarray data classificationAccurate and ScalableSystem for Automatic…Accurate and Scalable System for Automatic Detection of Malignant MelanomaFC-MST: Featurecorrelation maximum…FC-MST: Feature correlation maximum spanning tree for multimedia concept classificationA survey on featuredrift adaptation…A survey on feature drift adaptation: Definition, benchmark, challenges and future directionsA Bayesian Network Modelfor Predicting…A Bayesian Network Model for Predicting Post-stroke Outcomes With Available Risk FactorsAn efficient featuregeneration approach…An efficient feature generation approach based on deep learning and feature selection techniques for traffic classificationFeature Selection forHigh-Dimensional Data…Feature Selection for High-Dimensional Data: A Fast Correlation-Based Filter Solution過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。