SearcharxivSearch

arXiv subjects

Jyotishka Ray Choudhury

Publications and source records attributed to Jyotishka Ray Choudhury.

5 recordsLinked to original sources

Saddlepoint approximations for plug-in resampling

Resampling-based procedures can improve on normal approximations in sparse, large-scale testing problems, but their computational cost can be prohibitive. We recognize that several existing procedures belong to a faster plug-in resampling subclass, fixing fitted nuisance parameters during resampling. When the resampled statistic is a sum of conditionally independent terms, the saddlepoint approximation (SPA) for the resampling $p$-value offers further acceleration, replacing resampling with an analytical tail approximation. However, standard Edgeworth-based approximation-error bounds impose regularity conditions that are hard to verify for plug-in resampling laws. We use an alternative approach to establish a finite-sample relative-error bound for the Lugannani-Rice approximation under more tractable conditions, which we apply in two contexts. In statistical genetics, we identify response resampling procedures as the targets of existing SPAs and establish guarantees in a representative setting. In conditional independence testing, we introduce spaCRT, an SPA for the distilled conditional randomization test (dCRT), which has been applied successfully in biology. Our rates quantify the effects of sparsity and signal strength, with matching lower bounds in special cases. We additionally establish asymptotic Type-I error control of the corresponding plug-in resampling procedures under growing sparsity. In simulations and single-cell CRISPR data analysis, spaCRT closely approximates dCRT $p$-values and preserves its statistical performance while accelerating computation by up to 250-fold.

stat.ME

High-Dimensional Robust Change-Point Detection via Angular Kernel Statistics

We study nonparametric change-point detection for high-dimensional data in regimes where inference must be performed from small batches of observations. Our primary focus is the high-dimensional, low sample size (HDLSS) regime, where the sequence length is fixed while the ambient dimension diverges. We propose a dimension-averaged angular kernel scan framework for detecting marginal distributional shifts. The statistic aggregates bounded one-dimensional angular discrepancies across coordinates, yielding a fully nonparametric, hyperparameter-free, and moment-agnostic estimator that remains well-defined without specifying, estimating, or assuming finite marginal moments; for example, under heavy-tailed or contaminated distributions. For the offline single-change problem, we derive an exact population mean factorization into a universal deterministic shape function and a scalar signal factor, and characterize the exact null covariance structure up to a scalar variance factor, both valid for any fixed sample size and dimension. We also establish an HDLSS multivariate central limit theorem under cross-coordinate strong mixing which leads to a variance-calibrated asymptotically distribution-free test, asymptotic type-I error control, and lower bounds on power and localization accuracy. We further extend the offline procedure to a fixed-window sequential monitoring procedure for high-dimensional streaming data, and obtain ARL calibration and worst-case Pollak EDD bounds. Simulation studies demonstrate that the proposed method can accurately detect and localize changes in many challenging HDLSS and streaming high-dimensional settings where moment-based or hyperparameter-sensitive procedures may be extremely unstable or inaccurate.

stat.ME

The saddlepoint approximation for averages of conditionally independent random variables

Motivated by the application of saddlepoint approximations to resampling-based statistical tests, we prove that the Lugannani-Rice formula has vanishing relative error when applied to approximate conditional tail probabilities of averages of conditionally independent random variables. In a departure from existing work, this result is valid under only sub-exponential assumptions on the summands, and does not require any assumptions on their smoothness or lattice structure. The derived saddlepoint approximation result can be directly applied to resampling-based hypothesis tests, including bootstrap, sign-flipping and conditional randomization tests. We exemplify this by providing the first rigorous justification of a saddlepoint approximation for the sign-flipping test of symmetry about the origin, initially proposed in 1955. On the way to our main result, we establish a conditional Berry-Esseen inequality for sums of conditionally independent random variables, which may be of independent interest.

math.ST

Robust Classification of High-Dimensional Data using Data-Adaptive Energy Distance

Classification of high-dimensional low sample size (HDLSS) data poses a challenge in a variety of real-world situations, such as gene expression studies, cancer research, and medical imaging. This article presents the development and analysis of some classifiers that are specifically designed for HDLSS data. These classifiers are free of tuning parameters and are robust, in the sense that they are devoid of any moment conditions of the underlying data distributions. It is shown that they yield perfect classification in the HDLSS asymptotic regime, under some fairly general conditions. The comparative performance of the proposed classifiers is also investigated. Our theoretical results are supported by extensive simulation studies and real data analysis, which demonstrate promising advantages of the proposed classification techniques over several widely recognized methods.

stat.ML

Dirichlet Process-based Robust Clustering using the Median-of-Means Estimator

Clustering stands as one of the most prominent challenges in unsupervised machine learning. Among centroid-based methods, the classic $k$-means algorithm, based on Lloyd's heuristic, is widely used. Nonetheless, it is a well-known fact that $k$-means and its variants face several challenges, including heavy reliance on initial cluster centroids, susceptibility to converging into local minima of the objective function, and sensitivity to outliers and noise in the data. When data contains noise or outliers, the Median-of-Means (MoM) estimator offers a robust alternative for stabilizing centroid-based methods. On a different note, another limitation in many commonly used clustering methods is the need to specify the number of clusters beforehand. Model-based approaches, such as Bayesian nonparametric models, address this issue by incorporating infinite mixture models, which eliminate the requirement for predefined cluster counts. Motivated by these facts, in this article, we propose an efficient and automatic clustering technique by integrating the strengths of model-based and centroid-based methodologies. Our method mitigates the effect of noise on the quality of clustering; while at the same time, estimates the number of clusters. Statistical guarantees on an upper bound of clustering error, and rigorous assessment through simulated and real datasets, suggest the advantages of our proposed method over existing state-of-the-art clustering algorithms.

stat.ML