An information-theoretic distance for k-NN classification: Enhancing mixed-type data analysis with Jensen-Shannon divergence

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Distance-based classifiers such as k-nearest neighbors (KNN) often struggle with categorical variables due to the absence of a meaningful distance metric. Simple overlap or one-hot encodings are commonly used but are insufficient. They fail to capture differences in the underlying class-conditional distributions, especially for attributes with many categories. We address this problem by proposing an information-theoretic mixed-type dis tance for KNN classification: the square root of the Jensen-Shannon divergence (JSD) for categorical attributes is combined with Euclidean distance on normalized numeric attributes, and each feature is further scaled by its mutual information (MI) with the class label. This construction supplies a principled distance for categori cal levels, aligns heterogeneous scales, and enhances discriminative power. Across 27 benchmark datasets, our JSD-based method achieved the highest mean accuracy of 0.8433, outperforming standard baselines (Dummy and numeric-only KNN) and established heterogeneous metrics (HEOM and HVDM). These results demonstrate that an information-theoretic, distribution-aware distance provides an effective and interpretable solution to mixed-type KNN classification.

키워드

Mixed-type data; k-nearest neighbors; Jensen-Shannon divergence; Metric learning; Mutual information weighting; ALGORITHM
제목
An information-theoretic distance for k-NN classification: Enhancing mixed-type data analysis with Jensen-Shannon divergence
저자
Kim, Donggyu; Cho, HyungJun
DOI
10.1016/j.neucom.2026.134650
발행일
2026-11-14
유형
Article
저널명
Neurocomputing
권
702