SearcharxivSearch

arXiv subjects

Prajamitra Bhuyan

Publications and source records attributed to Prajamitra Bhuyan.

12 recordsLinked to original sources

BLOC: A Global Optimization Framework for Sparse Covariance Estimation with Non-Convex Penalties

We introduce BLOC (Black-box Optimization over Correlation matrices), a general framework for sparse covariance estimation with non-convex penalties. BLOC operates on the manifold of correlation matrices and reparameterizes it via an angular Cholesky mapping, transforming the positive-definite, unit-diagonal constraint into an unconstrained search over a Euclidean hyperrectangle. This enables gradient-free global optimization of diverse objectives, including non-differentiable or black-box losses, using a pattern search routine with adaptive coordinate polling, run-wise restarts to escape local minima, and leveraging up to $d(d-1)$ parallel threads when optimizing a $d$-dimensional correlation matrix. The method is penalty-agnostic and ensures that every iterate is a valid correlation matrix, from which covariance estimates are obtained. We establish convergence guarantees, including stationarity, probabilistic escape from poor local minima, and sublinear rates under smooth convex losses. From a statistical perspective, we prove consistency, convergence rates, and sparsistency for penalized correlation estimators under general conditions, extending sparse covariance theory beyond the Gaussian setting. Empirically, BLOC with nonconvex penalties such as SCAD and MCP outperforms leading estimators in both low- and high-dimensional regimes, achieving lower estimation error and improved sparsity recovery. A parallel implementation enhances scalability, and a proteomic network application demonstrates robust, positive-definite sparse covariance estimation.

stat.ME

Modeling Zero-Inflated Longitudinal Circular Data Using Bayesian Methods: Application to Ophthalmology

This paper introduces the modeling of circular data with excess zeros under a longitudinal framework, where the response is a circular variable and the covariates can be both linear and circular in nature. In the literature, various circular-circular and circular-linear regression models have been studied and applied to different real-world problems. However, there are no models for addressing zero-inflated circular observations in the context of longitudinal studies. Motivated by a real case study, a mixed-effects two-stage model based on the projected normal distribution is proposed to handle such issues. The interpretation of the model parameters is discussed and identifiability conditions are derived. A Bayesian methodology based on Gibbs sampling technique is developed for estimating the associated model parameters. Simulation results show that the proposed method outperforms its competitors in various situations. A real dataset on post-operative astigmatism is analyzed to demonstrate the practical implementation of the proposed methodology. The use of the proposed method facilitates effective decision-making for treatment choices and in the follow-up phases.

stat.ME

Multiple Testing of Partial Conjunction Hypotheses for Assessing Replicability Across Dependent Studies

Replicability is central to scientific progress, and the partial conjunction (PC) hypothesis testing framework provides an objective tool to quantify it across disciplines. Existing PC methods assume independent studies. Yet many modern applications, such as genome-wide association studies (GWAS) with sample overlap, violate this assumption, leading to dependence among study-specific summary statistics. Failure to account for this dependence can drastically inflate type I errors when combining inferences. We propose e-Filter, a powerful procedure grounded on the theory of e-values. It involves a filtering step that retains a set of the most promising PC hypotheses, and a selection step where PC hypotheses from the filtering step are marked as discoveries whenever their e-values exceed a selection threshold. We establish the validity of e-Filter for FWER and FDR control under unknown study dependence. A comprehensive simulation study demonstrates its excellent power gains over competing methods. We apply e-Filter to a GWAS replicability study to identify consistent genetic signals for low-density lipoprotein cholesterol (LDL-C). Here, the participating studies exhibit varying levels of sample overlap, rendering existing methods unsuitable for combining inferences. A subsequent pathway enrichment analysis shows that e-Filter replicated signals achieve stronger statistical enrichment on biologically relevant LDL-C pathways than competing approaches.

stat.ME

Estimating Prevalence of Post-war Health Disorders Using Capture-recapture Data

Effective surveillance on the long-term public health impact due to war and terrorist attacks remain limited. Such health issues are commonly under-reported, specifically for a large group of individuals. For this purpose, efficient estimation of the size of the population under the risk of physical and mental health hazards is of utmost necessity. In this context, multiple system estimation is a potential strategy that has recently been applied to quantify under-reported events allowing heterogeneity among the individuals and dependence between the sources of information. To model such complex phenomena, a novel trivariate Bernoulli model is developed, and an estimation methodology using Monte Carlo based EM algorithm is proposed. Simulation results show superiority of the performance of the proposed method over existing competitors and robustness under model mis-specifications. The method is applied to analyze real case studies on the Gulf War and 9/11 Terrorist Attack at World Trade Center, US. The results provide interesting insights that can assist in effective decision making and policy formulation for monitoring the health status of post-war survivors.

stat.AP

Causal Analysis at Extreme Quantiles with Application to London Traffic Flow Data

Transport engineers employ various interventions to enhance traffic-network performance. Quantifying the impacts of Cycle Superhighways is complicated due to the non-random assignment of such an intervention over the transport network. Treatment effects on asymmetric and heavy-tailed distributions are better reflected at extreme tails rather than at the median. We propose a novel method to estimate the treatment effect at extreme tails incorporating heavy-tailed features in the outcome distribution. The analysis of London transport data using the proposed method indicates that the extreme traffic flow increased substantially after Cycle Superhighways came into operation.

stat.ME

Estimation of Population Size with Heterogeneous Catchability and Behavioural Dependence: Applications to Air and Water Borne Disease Surveillance

Population size estimation based on the capture-recapture experiment is an interesting problem in various fields including epidemiology, criminology, demography, etc. In many real-life scenarios, there exists inherent heterogeneity among the individuals and dependency between capture and recapture attempts. A novel trivariate Bernoulli model is considered to incorporate these features, and the Bayesian estimation of the model parameters is suggested using data augmentation. Simulation results show robustness under model misspecification and the superiority of the performance of the proposed method over existing competitors. The method is applied to analyse real case studies on epidemiological surveillance. The results provide interesting insight on the heterogeneity and dependence involved in the capture-recapture mechanism. The methodology proposed can assist in effective decision-making and policy formulation.

stat.ME

On a Bivariate Copula for Modeling Negative Dependence: Application to New York Air Quality Data

In many practical scenarios, including finance, environmental sciences, system reliability, etc., it is often of interest to study the various notion of negative dependence among the observed variables. A new bivariate copula is proposed for modeling negative dependence between two random variables that complies with most of the popular notions of negative dependence reported in the literature. Specifically, the Spearman's rho and the Kendall's tau for the proposed copula have a simple one-parameter form with negative values in the full range. Some important ordering properties comparing the strength of negative dependence with respect to the parameter involved are considered. Simple examples of the corresponding bivariate distributions with popular marginals are presented. Application of the proposed copula is illustrated using a real data set on air quality in the New York City, USA.

stat.ME

Analysing the causal effect of London cycle superhighways on traffic congestion

Transport operators have a range of intervention options available to improve or enhance their networks. Such interventions are often made in the absence of sound evidence on resulting outcomes. Cycling superhighways were promoted as a sustainable and healthy travel mode, one of the aims of which was to reduce traffic congestion. Estimating the impacts that cycle superhighways have on congestion is complicated due to the non-random assignment of such intervention over the transport network. In this paper, we analyse the causal effect of cycle superhighways utilising pre-intervention and post-intervention information on traffic and road characteristics along with socio-economic factors. We propose a modeling framework based on the propensity score and outcome regression model. The method is also extended to the doubly robust set-up. Simulation results show the superiority of the performance of the proposed method over existing competitors. The method is applied to analyse a real dataset on the London transport network. The methodology proposed can assist in effective decision making to improve network performance.

stat.AP

On the Estimation of Population Size from a Dependent Triple Record System

Population size estimation based on capture-recapture experiment under triple record system is an interesting problem in various fields including epidemiology, population studies, etc. In many real life scenarios, there exists inherent dependency between capture and recapture attempts. We propose a novel model that successfully incorporates the possible dependency and the associated parameters possess nice interpretations. We provide estimation methodology for the population size and the associated model parameters based on maximum likelihood method. The proposed model is applied to analyze real data sets from public health and census coverage evaluation study. The performance of the proposed estimate is evaluated through extensive simulation study and the results are compared with the existing competitors. The results exhibit superiority of the proposed model over the existing competitors both in real data analysis and simulation study.

stat.ME

Two-stage Circular-circular Regression with Zero-inflation: Application to Medical Sciences

This paper considers the modeling of zero-inflated circular measurements concerning real case studies from medical sciences. Circular-circular regression models have been discussed in the statistical literature and illustrated with various real-life applications. However, there are no models to deal with zero-inflated response as well as a covariate simultaneously. The Mobius transformation based two-stage circular-circular regression model is proposed, and the Bayesian estimation of the model parameters is suggested using the MCMC algorithm. Simulation results show the superiority of the performance of the proposed method over the existing competitors. The method is applied to analyse real datasets on astigmatism due to cataract surgery and abnormal gait related to orthopaedic impairment. The methodology proposed can assist in efficient decision making during treatment or post-operative care.

stat.ME

Optimal Replacement Policy under Cumulative Damage Model and Strength Degradation with Applications

In many real-life scenarios, system failure depends on dynamic stress-strength interference, where strength degrades and stress accumulates concurrently over time. In this paper, we consider the problem of finding an optimal replacement strategy that balances the cost of replacement with the cost of failure and results in a minimum expected cost per unit time under cumulative damage model with strength degradation. The existing recommendations are applicable only under restricted distributional assumptions and/or with fixed strength. As theoretical evaluation of the expected cost per unit time turns out to be very complicated, a simulation-based algorithm is proposed to evaluate the expected cost rate and find the optimal replacement strategy. The proposed method is easy to implement having wider domain of application. For illustration, the proposed method is applied to real case studies on mailbox and cell-phone battery experiments.

stat.AP

On the estimation of population size from a post-stratified two sample capture-recapture data under dependence

Population size estimation based on two sample capture-recapture type experiment is an interesting problem in various fields including epidemiology, pubic health, population studies, etc. The Lincoln-Petersen estimate is popularly used under the assumption that capture and recapture status of each individual is independent. However, in many real life scenarios, there is an inherent dependency between capture and recapture attempts which is not well-studied in the literature of the dual system or two sample capture-recapture method. In this article, we propose a novel model that successfully incorporates the possible causal dependency and provide corresponding estimation methodologies for the associated model parameters based on post-stratified two sample capture-recapture data. The superiority of the performance of the proposed model over the existing competitors is established through an extensive simulation study. The method is illustrated through analysis of some real data sets.

stat.ME