Searcharxiv⌕ Search

arXiv subjects

Swapnaneel Bhattacharyya

Publications and source records attributed to Swapnaneel Bhattacharyya.

6 recordsLinked to original sources

Change detection with conformal martingales: new optimal constructions, and suboptimality of existing methods

We study distribution-free sequential changepoint detection for independent observations with unknown and unrestricted pre- and post-change laws. We build on the conformal test martingales and associated e-detectors of Vovk(2021), which control the probability of false alarm (PFA) and the average run length (ARL) respectively. The majority of these works focus on validity, with statistical efficiency usually left for simulations. We develop a comprehensive theory of how conformal p-values behave under non-exchangeable data with a changepoint at an unknown time $T$. We use this to analyze the post-change growth and resulting detection delay of conformal martingale methods, and prove that the standard existing methods are suboptimal for PFA and ARL control, and can lead to delays that are $Ω(T)$ and $Ω(\sqrt{\text{ARL}})$ respectively. We propose different conformal e-processes and e-detectors that are provably minimax optimal, with delays $Θ(\log T)$ and $Θ(\log \text{ARL})$ respectively, and have much shorter delays in simulations.

math.ST↗

Theoretical guarantees for change localization using conformal p-values

Changepoint localization aims to provide confidence sets for a changepoint (if one exists). Existing methods either relying on strong parametric assumptions or providing only asymptotic guarantees or focusing on a particular kind of change(e.g., change in the mean) rather than the entire distributional change. A method (possibly the first) to achieve distribution-free changepoint localization with finite-sample validity was recently introduced by \cite{dandapanthula2025conformal}. However, while they proved finite sample coverage, there was no analysis of set size. In this work, we provide rigorous theoretical guarantees for their algorithm. We also show the consistency of a point estimator for change, and derive its convergence rate without distributional assumptions. Along that line, we also construct a distribution-free consistent test to assess whether a particular time point is a changepoint or not. Thus, our work provides unified distribution-free guarantees for changepoint detection, localization, and testing. In addition, we present various finite sample and asymptotic properties of the conformal $p$-value in the distribution change setup, which provides a theoretical foundation for many applications of the conformal $p$-value. As an application of these properties, we construct distribution-free consistent tests for exchangeability against distribution-change alternatives and a new, computationally tractable method of optimizing the powers of conformal tests. We run detailed simulation studies to corroborate the performance of our methods and theoretical results. Together, our contributions offer a comprehensive and theoretically principled approach to distribution-free changepoint inference, broadening both the scope and credibility of conformal methods in modern changepoint analysis.

math.ST↗

Testing Equality of Medians for Multiple Samples

In this paper, we construct a consistent non-parametric test for testing the equality of population medians for different samples when the observations in each sample are independent and identically distributed. This test can be further used to test the equality of unknown location parameters for different samples. The method discussed in this paper can be extended to any quantile level instead of the median. We present the theoretical results and also demonstrate the performance of this test through simulation studies.

stat.ME↗

Application of Random Matrix Theory in High-Dimensional Statistics

This review article provides an overview of random matrix theory (RMT) with a focus on its growing impact on the formulation and inference of statistical models and methodologies. Emphasizing applications within high-dimensional statistics, we explore key theoretical results from RMT and their role in addressing challenges associated with high-dimensional data. The discussion highlights how advances in RMT have significantly influenced the development of statistical methods, particularly in areas such as covariance matrix inference, principal component analysis (PCA), signal processing, and changepoint detection, demonstrating the close interplay between theory and practice in modern high-dimensional statistical inference.

stat.ME↗

Analysis of Pleiotropy for Testosterone and Lipid Profiles in Males and Females

In modern scientific studies, it is often imperative to determine whether a set of phenotypes is affected by a single factor. If such an influence is identified, it becomes essential to discern whether this effect is contingent upon categories such as sex or age group, and importantly, to understand whether this dependence is rooted in purely non-environmental reasons. The exploration of such dependencies often involves studying pleiotropy, a phenomenon wherein a single genetic locus impacts multiple traits. This heightened interest in uncovering dependencies by pleiotropy is fueled by the growing accessibility of summary statistics from genome-wide association studies (GWAS) and the establishment of thoroughly phenotyped sample collections. This advancement enables a systematic and comprehensive exploration of the genetic connections among various traits and diseases. additive genetic correlation illuminates the genetic connection between two traits, providing valuable insights into the shared biological pathways and underlying causal relationships between them. In this paper, we present a novel method to analyze such dependencies by studying additive genetic correlations between pairs of traits under consideration. Subsequently, we employ matrix comparison techniques to discern and elucidate sex-specific or age-group-specific associations, contributing to a deeper understanding of the nuanced dependencies within the studied traits. Our proposed method is computationally handy and requires only GWAS summary statistics. We validate our method by applying it to the UK Biobank data and present the results.

stat.ME↗

A Statistical Approach to Ecological Modeling by a New Similarity Index

Similarity index is an important scientific tool frequently used to determine whether different pairs of entities are similar with respect to some prefixed characteristics. Some standard measures of similarity index include Jaccard index, Sørensen-Dice index, and Simpson's index. Recently, a better index ($\hatα$) for the co-occurrence and/or similarity has been developed, and this measure really outperforms and gives theoretically supported reasonable predictions. However, the measure $\hatα$ is not data dependent. In this article we propose a new measure of similarity which depends strongly on the data before introducing randomness in prevalence. Then, we propose a new method of randomization which changes the whole pattern of results. Before randomization our measure is similar to the Jaccard index, while after randomization it is close to $\hatα$. We consider the popular ecological dataset from the Tuscan Archipelago, Italy; and compare the performance of the proposed index to other measures. Since our proposed index is data dependent, it has some interesting properties which we illustrate in this article through numerical studies.

stat.ME↗