SearcharxivSearch

arXiv subjects

Dan Zhuang

Publications and source records attributed to Dan Zhuang.

5 recordsLinked to original sources

Change-Point Detection for Heterogeneous High-Dimensional Functional Time Series

High-dimensional functional panels consist of temporally ordered curves observed across many subjects and naturally exhibit heterogeneous structural changes. Under sparse subject-level break signals or opposite-signed shifts, traditional mean-aggregated CUSUM procedures may suffer noticeable power loss due to signal attenuation or cancellation induced by cross-sectional averaging. We propose a novel Energy--PE statistic, which combines subject-wise squared CUSUM energy aggregation with a generalized power-enhancement component. The energy aggregation preserves subject-level evidence under sign-heterogeneous changes, while the power-enhancement component improves sensitivity to sparse weak break signals. Under regularity conditions, we establish the asymptotic behavior of the proposed statistic. We further incorporate a latent group structure and an information-criterion-based clustering algorithm to estimate the unknown group number and membership for heterogeneous break points. Numerical studies and an intraday stock application demonstrate that Energy--PE controls size, improves power under sparse and sign-heterogeneous alternatives, and yields interpretable post-test summaries.

stat.ME

Semiparametric Elliptical Mixture Clustering for High-Dimensional Data

Clustering high-dimensional data is especially challenging when cluster distributions are heavy tailed and only approximately elliptical. Existing high-dimensional methods are largely built for Gaussian or other light-tailed models, whereas classical robust elliptical procedures are mostly low dimensional or rely on fully parametric radial families. We propose a semiparametric elliptical mixture clustering framework with cluster-specific centers, an unknown common radial generator, and a common sparse precision-shape matrix, together with a data-driven rule for selecting the number of clusters. A generalized expectation-maximization (GEM) algorithm is developed by combining transformed-radius estimation of the radial generator, radial-score center updates, and a Tyler-POET-GLASSO update for the common precision-shape matrix. The method avoids specifying a parametric radial family and remains computationally feasible in high dimensions. We establish high-dimensional consistency for the estimated model components and the excess misclustering error. Simulation studies and a handwritten-digit application demonstrate the competitive performance and robustness of the proposed method, particularly in heavy-tailed elliptical settings.

stat.ME

Sparse $K$-spatial-median clustering for high-dimensional data

We propose a robust clustering framework for high-dimensional data with heavy tails and a large fraction of irrelevant variables. The method replaces the mean updates of Lloyd's $K$-means with \emph{spatial medians} to enhance robustness. For the assignment step, it admits either a Euclidean rule for computational simplicity or a robust Mahalanobis-type metric constructed from the spatial sign covariance matrix to account for heterogeneous scales and feature dependence. To handle the $p \gg n$ regime, we further introduce a simple \emph{hard feature-exclusion} mechanism that removes weakly separating dimensions based on across-center dispersion, with the exclusion threshold selected automatically via a permutation-based Gap criterion. Simulation studies under correlated Gaussian and multivariate $t$ models demonstrate that the proposed approach provides competitive clustering accuracy and improved stability relative to $K$-means and sparse $K$-means baselines.

stat.ME

Adaptive Test for High Dimensional Quantile Regression

Testing high-dimensional quantile regression coefficients is crucial, as tail quantiles often reveal more than the mean in many practical applications. Nevertheless, the sparsity pattern of the alternative hypothesis is typically unknown in practice, posing a major challenge. To address this, we propose an adaptive test that remains powerful across both sparse and dense alternatives.We first establish the asymptotic independence between the max-type test statistic proposed by \citet{tang2022conditional} and the sum-type test statistic introduced by \citet{chen2024hypothesis}. Building on this result, we propose a Cauchy combination test that effectively integrates the strengths of both statistics and achieves robust performance across a wide range of sparsity levels. Simulation studies and real data applications demonstrate that our proposed procedure outperforms existing methods in terms of both size control and power.

stat.ME

Spatial Sign based Direct Sparse Linear Discriminant Analysis for High Dimensional Data

This paper investigates the robust linear discriminant analysis (LDA) problem with elliptical distributions in high-dimensional data. We propose a robust classification method, named SSLDA, that is intended to withstand heavy-tailed distributions. We demonstrate that SSLDA achieves an optimal convergence rate in terms of both misclassification rate and estimate error. Our theoretical results are further confirmed by extensive numerical experiments on both simulated and real datasets. Compared with current approaches, the SSLDA method offers superior improved finite sample performance and notable robustness against heavy-tailed distributions.

stat.ME