Searcharxiv⌕ Search

arXiv subjects

Bhargab Chattopadhyay

Publications and source records attributed to Bhargab Chattopadhyay.

5 recordsLinked to original sources

A Three-Stage PCA Procedure for Sequentially Arriving High-Dimensional Data

We develop a three-stage adaptive procedure for principal component analysis (PCA) when high-dimensional observations are collected sequentially and additional sampling incurs a cost. The procedure balances PCA compression loss against sampling cost while selecting the retained dimension through a prescribed explained-variance criterion. Starting from a pilot sample, an intermediate stage updates the PCA quantities before determining the final sample size, thereby avoiding reliance on unknown population eigenvalues. Under suitable regularity conditions, we establish both first- and second-order efficiency relative to the population oracle. Comparison with the corresponding two-stage rule shows that the additional recalibration yields sharper second-order control and reduces the influence of the pilot stage on the final sampling decision. The theory allows the ambient dimension to exceed the sample size under appropriate covariance and spectral conditions. Simulation studies demonstrate the strong finite-sample performance of the procedure across increasing dimensions and several dense covariance structures. As a real-data application, we conduct a retrospective study of gene-expression data from 32 cancer-type cohorts in The Cancer Genome Atlas, illustrating both cost-effective early stopping and settings in which additional observations are recommended.

stat.ME↗

Extreme Population Selection under Multistage Sampling design With Applications

We study the problem of selecting the extreme (best or worst) population from among $K(\geq 2)$ populations, under the assumption that the extreme population is sufficiently separated from the nearest population. The selection is based on an appropriate measure, which may vary across different application domains. Since the actual value of the measure is unknown, we obtain its estimator using the generalized method of moments under a multistage sampling design. Using this estimator, we propose two sequential procedures, namely online algorithm and multi armed bandit based algorithm. Under suitable regularity conditions and without imposing parametric assumptions on the underlying distributions, both algorithms correctly identify the extreme population with a desired level of confidence. We illustrate the proposed algorithms through applications in econometrics and genetics. In the econometric application, the extreme population is selected using the Gini index as a measure of inequality and the performance of the proposed procedures is assessed through extensive Monte Carlo simulation studies conducted under various distributional settings. In the genetics application, the worst population is identified using a measure derived from the tumor mutation burden (TMB) score and the practical applicability of the proposed algorithms is demonstrated using the Memorial Sloan Kettering-IMPACT 50000 clinical sequencing cohort. Further, we use the proposed framework to identify an anomalous population, provided such a population exists.

stat.ME↗

Asymptotically Optimal Sequential Confidence Interval for the Gini Index Under Complex Household Survey Design with Sub-Stratification

We examine the optimality properties of the Gini index estimator under complex survey design involving stratification, clustering, and sub-stratification. While Darku et al. (Econometrics, 26, 2020) considered only stratification and clustering and did not provide theoretical guarantees, this study addresses these limitations by proposing two procedures - a purely sequential method and a two-stage method. Under suitable regularity conditions, we establish uniform continuity in probability for the proposed estimator, thereby contributing to the development of random central limit theorems under sequential sampling frameworks. Furthermore, we show that the resulting procedures satisfy both asymptotic first-order efficiency and asymptotic consistency. Simulation results demonstrate that the proposed procedures achieve the desired optimality properties across diverse settings. The practical utility of the methodology is further illustrated through an empirical application using data collected by the National Sample Survey agency of India

stat.ME↗

Minimum Risk Point Estimation of Gini Index

This paper develops a theory and methodology for estimation of Gini index such that both cost of sampling and estimation error are minimum. Methods in which sample size is fixed in advance, cannot minimize estimation error and sampling cost at the same time. In this article, a purely sequential procedure is proposed which provides an estimate of the sample size required to achieve a sufficiently smaller estimation error and lower sampling cost. Characteristics of the purely sequential procedure are examined and asymptotic optimality properties are proved without assuming any specific distribution of the data. Performance of our method is examined through extensive simulation study.

stat.ME↗

Estimation of Gini Index within Pre-Specied Error Bound

Gini index is a widely used measure of economic inequality. This article develops a general theory for constructing a confidence interval for Gini index with a specified confidence coefficient and a specified width. Fixed sample size methods cannot simultaneously achieve both the specified confidence coefficient and specified width. We develop a purely sequential procedure for interval estimation of Gini index with a specified confidence coefficient and a fixed margin of error. Optimality properties of the proposed method, namely first order asymptotic efficiency and asymptotic consistency are proved. All theoretical results are derived without assuming any specific distribution of the data.

stat.ME↗