Selection of Relevant Features in Machine Learning.

In this paper, we review the problem of selecting relevant features for use in machine learning. We describe this problem in terms of heuristic search through a space of feature sets, and we identify four dimensions along which approaches to the problem can vary. We consider recent work on feature selection in terms of this framework, then close with some challenges for future work in the area. 1. The Problem of Irrelevant Features The selection of relevant features, and the elimination of irrelevant ones, is a central problem in machine learning. Before an induction algorithm can move beyond the training data to make predictions about novel test cases, it must decide which attributes to use in these predictions and which to ignore. Intuitively, one would like the learner to use only those attributes that are `relevant' to the target concept. There have been a few attempts to define `relevance' in the context of machine learning, as John, Kohavi, and Pfleger (1994) have noted in their...

Selection of Relevant Features in Machine Learning. | Litlas