Searcharxiv⌕ Search

arXiv subjects

Srijan Chattopadhyay

Publications and source records attributed to Srijan Chattopadhyay.

5 recordsLinked to original sources

On the optimality of antithetic randomization for cross-validation

In the classical normal means problem, independent train--test folds can be constructed by perturbing the data with normal randomization. Averaging over $K$ such folds yields a cross-validation estimator whose bias depends on the marginal distribution of the randomization variables, while its variance depends on their joint distribution. This raises the questions: which joint law is optimal, and how to construct the corresponding randomization scheme? We show that: (i) for smooth estimators, antithetic randomization with pairwise correlation $ρ=-1/(K-1)$ is necessary and sufficient for the reducible variance due to randomization to remain bounded as the bias vanishes; (ii) a general construction yields a class of antithetic schemes, within which the jointly normal scheme is minimax optimal; and (iii) for non-smooth estimators with finitely many jump discontinuities, antithetic randomization improves the asymptotic rate of the reducible variance, while a simple control variate restores bounded variance when the discontinuities are known.

math.ST↗

Inference for quantile-parametrized families via CDF confidence bands

Quantile-based distribution families are an important subclass of parametric families, capable of exhibiting a wide range of behaviors using very few parameters. These parametric models present significant challenges for classical methods, since the CDF and density do not have a closed-form expression. Furthermore, approximate maximum likelihood estimation and related procedures may yield non-$\sqrt{n}$ and non-normal asymptotics over regions of the parameter space, making bootstrap and resampling techniques unreliable. We develop a novel inference framework that constructs confidence sets by inverting distribution-free confidence bands for the empirical CDF through the known quantile function. Our proposed inference procedure provides a principled and assumption-lean alternative in this setting, requiring no distributional assumptions beyond the parametric model specification and avoiding the computational and theoretical difficulties associated with likelihood-based methods for these complex parametric families. We demonstrate our framework on Tukey Lambda and generalized Lambda distributions, evaluate its performance through simulation studies, and illustrate its practical utility with an application to both a small-sample dataset (Twin Study) and a large-sample dataset (Spanish household incomes).

stat.ME↗

Application of Random Matrix Theory in High-Dimensional Statistics

This review article provides an overview of random matrix theory (RMT) with a focus on its growing impact on the formulation and inference of statistical models and methodologies. Emphasizing applications within high-dimensional statistics, we explore key theoretical results from RMT and their role in addressing challenges associated with high-dimensional data. The discussion highlights how advances in RMT have significantly influenced the development of statistical methods, particularly in areas such as covariance matrix inference, principal component analysis (PCA), signal processing, and changepoint detection, demonstrating the close interplay between theory and practice in modern high-dimensional statistical inference.

stat.ME↗

Analysis of Pleiotropy for Testosterone and Lipid Profiles in Males and Females

In modern scientific studies, it is often imperative to determine whether a set of phenotypes is affected by a single factor. If such an influence is identified, it becomes essential to discern whether this effect is contingent upon categories such as sex or age group, and importantly, to understand whether this dependence is rooted in purely non-environmental reasons. The exploration of such dependencies often involves studying pleiotropy, a phenomenon wherein a single genetic locus impacts multiple traits. This heightened interest in uncovering dependencies by pleiotropy is fueled by the growing accessibility of summary statistics from genome-wide association studies (GWAS) and the establishment of thoroughly phenotyped sample collections. This advancement enables a systematic and comprehensive exploration of the genetic connections among various traits and diseases. additive genetic correlation illuminates the genetic connection between two traits, providing valuable insights into the shared biological pathways and underlying causal relationships between them. In this paper, we present a novel method to analyze such dependencies by studying additive genetic correlations between pairs of traits under consideration. Subsequently, we employ matrix comparison techniques to discern and elucidate sex-specific or age-group-specific associations, contributing to a deeper understanding of the nuanced dependencies within the studied traits. Our proposed method is computationally handy and requires only GWAS summary statistics. We validate our method by applying it to the UK Biobank data and present the results.

stat.ME↗

A Statistical Approach to Ecological Modeling by a New Similarity Index

Similarity index is an important scientific tool frequently used to determine whether different pairs of entities are similar with respect to some prefixed characteristics. Some standard measures of similarity index include Jaccard index, Sørensen-Dice index, and Simpson's index. Recently, a better index ($\hatα$) for the co-occurrence and/or similarity has been developed, and this measure really outperforms and gives theoretically supported reasonable predictions. However, the measure $\hatα$ is not data dependent. In this article we propose a new measure of similarity which depends strongly on the data before introducing randomness in prevalence. Then, we propose a new method of randomization which changes the whole pattern of results. Before randomization our measure is similar to the Jaccard index, while after randomization it is close to $\hatα$. We consider the popular ecological dataset from the Tuscan Archipelago, Italy; and compare the performance of the proposed index to other measures. Since our proposed index is data dependent, it has some interesting properties which we illustrate in this article through numerical studies.

stat.ME↗