Conditional Likelihood Maximisation: A Unifying Framework for Information Theoretic Feature Selection

We present a unifying framework for information theoretic feature selection, bringing almost two decades of research on heuristic filter criteria under a single theoretical interpretation. This is in response to the question: “what are the implicit statistical assumptions of feature selection criteria based on mutual information?”. To answer this, we adopt a different strategy than is usual in the feature selection literature—instead of trying to define a criterion, we derive one, directly from a clearly specified objective function: the conditional likelihood of the training labels. While many hand-designed heuristic criteria try to optimize a definition of feature ‘relevancy ’ and ‘redundancy’, our approach leads to a probabilistic framework which naturally incorporates these concepts. As a result we can unify the numerous criteria published over the last two decades, and show them to be low-order approximations to the exact (but intractable) optimisation problem. The primary contribution is to show that common heuristics for information based feature selection (including Markov Blanket algorithms as a special case) are approximate iterative maximisers of the conditional likelihood. A large empirical study provides strong evidence to favour certain classes of criteria, in particular those that balance the relative size of the relevancy/redundancy terms. Overall we conclude that the JMI criterion (Yang and Moody, 1999; Meyer et al., 2008) provides the best tradeoff in terms of accuracy, stability, and flexibility with small data samples.

Feature selection andfeature extraction for…Feature selection and feature extraction for text categorizationToward Optimal FeatureSelectionToward Optimal Feature SelectionAlgorithms for LargeScale Markov Blanket…Algorithms for Large Scale Markov Blanket DiscoveryFast Binary FeatureSelection with…Fast Binary Feature Selection with Conditional Mutual InformationEfficient FeatureSelection via Analysis…Efficient Feature Selection via Analysis of Relevance and RedundancyFeature Selection Basedon Mutual Information…Feature Selection Based on Mutual Information: Criteria of Max-Dependency, Max-Relevance, and Min-RedundancyConditional InfomaxLearning: An Integrated…Conditional Infomax Learning: An Integrated Framework for Feature Extraction and FusionOn the Use of VariableComplementarity for…On the Use of Variable Complementarity for Feature Selection in Cancer ClassificationInformation-TheoreticFeature Selection in…Information-Theoretic Feature Selection in Microarray Data Using Variable ComplementarityStable feature selectionvia dense feature groupsStable feature selection via dense feature groupsA New Perspective forInformation Theoretic…A New Perspective for Information Theoretic Feature SelectionConditional MutualInformation-Based…Conditional Mutual Information-Based Feature Selection Analyzing for Synergy and RedundancyParallel FeatureSelection Inspired by…Parallel Feature Selection Inspired by Group TestingBooster in HighDimensional Data…Booster in High Dimensional Data ClassificationA New Approach forFeature Selection from…A New Approach for Feature Selection from Microarray Data Based on Mutual InformationMarkov Blanket FeatureSelection Using…Markov Blanket Feature Selection Using Representative SetsFeature selection byoptimizing a lower boun…Feature selection by optimizing a lower bound of conditional mutual informationSpeeding up joint mutualinformation feature…Speeding up joint mutual information feature selection with an optimization heuristicSupervised InfiniteFeature SelectionSupervised Infinite Feature SelectionFeature Selection withConditional Mutual…Feature Selection with Conditional Mutual Information Considering Feature InteractionA semi-parallelframework for greedy…A semi-parallel framework for greedy information-theoretic feature selectionWatermelon: a NovelFeature Selection Metho…Watermelon: a Novel Feature Selection Method Based on Bayes Error Rate Estimation and a New Interpretation of Feature Relevance and RedundancyA conditional-weightjoint relevance metric…A conditional-weight joint relevance metric for feature relevancy termMachine learningintegrated ensemble of…Machine learning integrated ensemble of feature selection methods followed by survival analysis for predicting breast cancer subtype specific miRNA biomarkersConditional LikelihoodMaximisation: A Unifyin…Conditional Likelihood Maximisation: A Unifying Framework for Information Theoretic Feature SelectionEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.