arXiv · 2206.10866
Nearest Neighbor Classification based on Imbalanced Data: A Statistical Approach
Abstract
When the competing classes in a classification problem are not of comparable size, many popular classifiers exhibit a bias towards larger classes, and the nearest neighbor classifier is no exception. To take care of this problem, we develop a statistical method for nearest neighbor classification based on such imbalanced data sets. First, we construct a classifier for the binary classification problem and then extend it for classification problems involving more than two classes. Unlike the existing oversampling or undersampling methods, our proposed classifiers do not need to generate any pseudo observations or remove any existing observations, hence the results are exactly reproducible. We establish the Bayes risk consistency of these classifiers under appropriate regularity conditions. Their superior performance over the existing methods is amply demonstrated by analyzing several benchmark data sets.
Explore related subjects
Keep this discovery
Anvit Garg, Anil K. Ghosh, Soham Sarkar. 2022-06-22. Nearest Neighbor Classification based on Imbalanced Data: A Statistical Approach. https://arxiv.org/abs/2206.10866
Cite the original work for its findings. Save a collection to share your selection of sources.