F1-Optimal Thresholding in the Multi-Label Setting

This paper investigates the properties of the widely-utilized F1 metric as used to evaluate the performance of multi-label classifiers. We show that given an un-informative binary classifier, F1-optimal thresholding is to predict all instances positive. More surprisingly, we prove a relationship between the optimal threshold and the best achievable F1 score over all thresholds. We demonstrate that macro-averaged F1, a commonly used multi-label performance metric, can conceal this extreme thresholding behavior. Finally, based on these properties of F1, we suggest average skill score as an alternative to macro-averaged F1 for multi-label classifi-cation. 1

F1-Optimal Thresholding in the Multi-Label Setting | Litlas