SearcharxivSearch

arXiv subjects

Hend Gabr

Publications and source records attributed to Hend Gabr.

2 recordsLinked to original sources

High-Dimensional BWDM: A Robust Nonparametric Clustering Validation Index for Large-Scale Data

Determining the appropriate number of clusters in unsupervised learning is a central problem in statistics and data science. Traditional validity indices such as Calinski-Harabasz, Silhouette, and Davies-Bouldin-depend on centroid-based distances and therefore degrade in high-dimensional or contaminated data. This paper proposes a new robust, nonparametric clustering validation framework, the High-Dimensional Between-Within Distance Median (HD-BWDM), which extends the recently introduced BWDM criterion to high-dimensional spaces. HD-BWDM integrates random projection and principal component analysis to mitigate the curse of dimensionality and applies trimmed clustering and medoid-based distances to ensure robustness against outliers. We derive theoretical results showing consistency and convergence under Johnson-Lindenstrauss embeddings. Extensive simulations demonstrate that HD-BWDM remains stable and interpretable under high-dimensional projections and contamination, providing a robust alternative to traditional centroid-based validation criteria. The proposed method provides a theoretically grounded, computationally efficient stopping rule for nonparametric clustering in modern high-dimensional applications.

stat.ML

Nonparametric Clustering Stopping Rule Based on Multivariate Median

This paper introduces a novel nonparametric criterion for determining the appropriate number of clusters, which is derived from the spatial median. The method is constructed to reconcile two competing objectives of cluster analysis: the preservation of internal homogeneity within clusters and the maximization of heterogeneity across clusters. To this end, the proposed algorithm optimizes the ratio of inter-cluster to intra-cluster variability, incorporating adjustments for both the sample size and the number of clusters. Unlike conventional techniques, the method is distribution-free and demonstrates robustness in the presence of outliers. Its properties were first examined through extensive simulation studies, followed by empirical evaluations on three applied datasets. To further assess comparative performance, the proposed procedure was benchmarked against 13 established algorithms for cluster number determination. In 11 of these comparisons, the proposed criterion exhibited superior performance, thereby underscoring its utility as a reliable and rigorous alternative for multivariate clustering applications.

stat.CO