On the Surprising Behavior of Distance Metrics in High Dimensional Spaces

Abstract. In recent years, the eect of the curse of high dimensionality has been studied in great detail on several problems such as clustering, nearest neighbor search, and indexing. In high dimensional space the data becomes sparse, and traditional indexing and algorithmic techniques fail from a eciency and/or eectiveness perspective. Recent research results show that in high dimensional space, the concept of proximity, distance or nearest neighbor may not even be qualitatively meaningful. In this paper, we view the dimensionality curse from the point of view of the di-stance metrics which are used to measure the similarity between objects. We specically examine the behavior of the commonly used Lk norm and show that the problem of meaningfulness in high dimensionality is sensitive to the value of k. For example, this means that the Manhat-tan distance metric (L1 norm) is consistently more preferable than the Euclidean distance metric (L2 norm) for high dimensional data mining applications. Using the intuition derived from our analysis, we introduce and examine a natural extension of the Lk norm to fractional distance metrics. We show that the fractional distance metric provides more mea-ningful results both from the theoretical and empirical perspective. The results show that fractional distance metrics can signicantly improve the eectiveness of standard clustering algorithms such as the k-means algorithm. 1

On the Surprising Behavior of Distance Metrics in High Dimensional Spaces | Litlas