SearcharxivSearch

arXiv subjects

Monika Bhattacharjee

Publications and source records attributed to Monika Bhattacharjee.

8 recordsLinked to original sources

Testing Microbiome Community Differences in High Dimensions: A Bootstrap Approach for Compositional Data

Understanding differences in microbial community structure is critical for uncovering risk factors and mechanisms underlying diseases such as colorectal cancer and preterm birth. Microbiome data present unique statistical challenges because they are compositional in nature, violating assumptions of many classical inference procedures. We propose an empirical bootstrap framework that enables robust hypothesis testing for equality of microbial community means across groups, including two-sample, paired, and multi-sample settings. The method accounts for the simplex structure of microbiome data and provides valid inference even in high-dimensional regimes. Through applications to two large-scale studies, fecal microbiota in colorectal adenoma and cancer patients, and vaginal microbiota in pregnancy with preterm birth outcomes-we demonstrate that our approach identifies clinically meaningful differences that conventional methods fail to detect, such as age-related differences in adenoma prevalence and race-associated disparities in vaginal microbiome composition. These results highlight the potential of resampling-based inference for advancing microbiome research, improving reproducibility, and uncovering clinically relevant microbial signatures.

stat.ME

Limiting Spectral Distribution of moderately large Kendall's correlation matrix and its application

We establish the limiting spectral distribution of Kendall's correlation matrices in the moderate high-dimensional regime where the dimension grows slower than the sample size. Our framework allows observations to be independent but not necessarily identically distributed, and accommodates both discrete and continuous data. Unlike existing results developed under i.i.d. observations, our approach remains valid under substantial distributional heterogeneity and also covers certain i.i.d. models beyond previously studied settings. Under mild symmetry and convergence conditions on some traces, we prove that the empirical spectral distribution of a properly centered and scaled Kendall's correlation matrix converges weakly almost surely to a deterministic, generally model-dependent limit. The analysis clarifies how distributional heterogeneity influences the limiting spectrum. As an application, we propose a graphical tool for detecting dependence among components in high-dimensional data and show that ignoring heterogeneity may lead to spurious detection of dependence.

math.ST

Joint convergence of sample cross-covariance matrices

Suppose $X$ and $Y$ are $p\times n$ matrices each with mean $0$, variance $1$ and where all moments of any order are uniformly bounded as $p,n \to \infty$. Moreover, the entries $(X_{ij}, Y_{ij})$ are independent across $i,j$ with a common correlation $ρ$. Let $C=n^{-1}XY^*$ be the sample cross-covariance matrix. We show that if $n, p\to \infty, p/n\to y\neq 0$, then $C$ converges in the algebraic sense and the limit moments depend only on $ρ$. Independent copies of such matrices with same $p$ but different $n$, say $\{n_l\}$, different correlations $\{ρ_l\}$, and different non-zero $y$'s, say $\{y_l\}$ also converge jointly and are asymptotically free. When $y=0$, the matrix $\sqrt{np^{-1}}(C-ρI_p)$ converges to an elliptic variable with parameter $ρ^2$. In particular, this elliptic variable is circular when $ρ=0$ and is semi-circular when $ρ=1$. If we take independent $C_l$, then the matrices $\{\sqrt{n_lp^{-1}}(C_l-ρ_l I_p)\}$ converge jointly and are also asymptotically free. As a consequence, the limiting spectral distribution of any symmetric matrix polynomial exists and has compact support.

math.PR

Asymptotic freeness of sample covariance matrices via embedding

We present an alternative proof of asymptotic freeness of independent sample covariance matrices, when the dimension and the sample size grow at the same rate, by embedding these matrices into Wigner matrices of a larger order and using asymptotic freeness of independent Wigner and deterministic matrices.

math.PR

Change Point Estimation in a Dynamic Stochastic Block Model

We consider the problem of estimating the location of a single change point in a dynamic stochastic block model. We propose two methods of estimating the change point, together with the model parameters. The first employs a least squares criterion function and takes into consideration the full structure of the stochastic block model and is evaluated at each point in time. Hence, as an intermediate step, it requires estimating the community structure based on a clustering algorithm at every time point. The second method comprises of the following two steps: in the first one, a least squares function is used and evaluated at each time point, but ignores the community structures and just considers a random graph generating mechanism exhibiting a change point. Once the change point is identified, in the second step, all network data before and after it are used together with a clustering algorithm to obtain the corresponding community structures and subsequently estimate the generating stochastic block model parameters. A comparison between these two methods is illustrated. Further, for both methods under their respective identifiability and certain additional regularity conditions, we establish rates of convergence and derive the asymptotic distributions of the change point estimators. The results are illustrated on synthetic data.

math.ST

Change Point Estimation in Panel Data with Temporal and Cross-sectional Dependence

We study the problem of detecting a common change point in large panel data based on a mean shift model, wherein the errors exhibit both temporal and cross-sectional dependence. A least squares based procedure is used to estimate the location of the change point. Further, we establish the convergence rate and obtain the asymptotic distribution of the least squares estimator. The form of the distribution is determined by the behavior of the norm difference of the means before and after the change point. Since the behavior of this norm difference is, a priori, unknown to the practitioner, we also develop a novel data driven adaptive procedure that provides valid confidence intervals for the common change point, without requiring any such knowledge. Numerical work based on synthetic data illustrates the performance of the estimator in finite samples under different settings of temporal and cross-sectional dependence, sample size and number of panels. Finally, we examine an application to financial stock data and discuss the identified change points.

math.ST

Common change point estimation in panel data from the least squares and maximum likelihood viewpoints

We establish the convergence rates and asymptotic distributions of the common break change-point estimators, obtained by least squares and maximum likelihood in panel data models and compare their asymptotic variances. Our model assumptions accommodate a variety of commonly encountered probability distributions and, in particular, models of particular interest in econometrics beyond the commonly analyzed Gaussian model, including the zero-inflated Poisson model for count data, and the probit and tobit models. We also provide novel results for time dependent data in the signal-plus-noise model, with emphasis on a wide array of noise processes, including Gaussian process, MA$(\infty)$ and $m$-dependent processes. The obtained results show that maximum likelihood estimation requires a stronger signal-to-noise model identifiability condition compared to its least squares counterpart. Finally, since there are three different asymptotic regimes that depend on the behavior of the norm difference of the model parameters before and after the change point, which cannot be realistically assumed to be known, we develop a novel data driven adaptive procedure that provides valid confidence intervals for the common break, without requiring a priori knowledge of the asymptotic regime the problem falls in.

math.ST

Large sample behaviour of high dimensional autocovariance matrices

The existence of limiting spectral distribution (LSD) of $\hatΓ_u+\hatΓ_u^*$, the symmetric sum of the sample autocovariance matrix $\hatΓ_u$ of order $u$, is known when the observations are from an infinite dimensional vector linear process with appropriate (strong) assumptions on the coefficient matrices. Under significantly weaker conditions, we prove, in a unified way, that the LSD of any symmetric polynomial in these matrices such as $\hatΓ_u+\hatΓ_u^*$, $\hatΓ_u\hatΓ_u^*$, $\hatΓ_u\hatΓ_u^*+\hatΓ_k\hatΓ_k^*$ exist. Our approach is through the more intuitive algebraic method of free probability in conjunction with the method of moments. Thus, we are able to provide a general description for the limits in terms of some freely independent variables. All the previous results follow as special cases. We suggest statistical uses of these LSD and related results in order determination and white noise testing.

math.ST