An Axiomatically Derived Measure for the Evaluation of Classification Algorithms
We address the general problem of finding suitable evalu-ation measures for classification systems. To this end, we adopt an axiomatic approach, i.e., we discuss a number of properties (“axioms”) that an evaluation measure for clas-sification should arguably satisfy. We start our analysis by addressing binary classification. We show that F1, nowadays considered a standard measure for the evaluation of binary classification systems, does not comply with a number of them, and should thus be considered unsatisfactory. We go on to discuss an alternative, simple evaluation measure for binary classification, that we call K, and show that it in-stead satisfies all the previously proposed axioms. We thus argue that researchers and practitioners should replace F1 with K in their everyday binary classification practice. We carry on our analysis by showing that K can be smoothly extended to deal with single-label multi-class classification, cost-sensitive classification, and ordinal classification. 1.
