SearcharxivSearch

arXiv subjects

Rajarshi Mukherjee

Publications and source records attributed to Rajarshi Mukherjee.

At least 19 recordsLinked to original sources

Improved Variance Estimation in Homoskedastic Nonparametric Random-Design Regression via a Two-Scale Approach

We study estimation of a constant conditional variance $σ^2$ in nonparametric regression with a $d$-dimensional random design. This is an important problem, and similar questions arise in causal inference. The regression function is $β_b$-Hölder smooth, the design density is $β_g$-Hölder smooth and bounded above and away from zero, and we consider the nonparametric regime $β_b>1$ and $d>4β_b$. Set $β_g^\star=β_b(1-4β_b/d)/\{1+2β_b/d+8(β_b/d)^2\}$. We give an estimator whose mean squared error is upper bounded by $Cn^{-4(β_b+1)/(d+4)}$ in the low-regularity regime when $0<β_g\leqβ_g^\star$. The low-regularity branch is based on a new two-scale construction: the covariate space is partitioned into cells, the local polynomial trend is projected out within each suitable cell, and the squared normalized contrast from one eligible close pair per cell is averaged across cells. In the high regularity regime when $β_g>β_g^\star$, a higher-order influence function estimator of Robins, Li, Tchetgen Tchetgen, and van der Vaart (2008) provides the rate $Cn^{-8β_b/(d+4β_b)}$. We also give an all-pairs ridge extension, which achieves the same two-scale rate, and evaluate the methods alongside a range of existing estimators in simulations.

math.ST

Nuisance Function Tuning and Sample Splitting for Optimally Estimating a Doubly Robust Functional

Estimators of doubly robust functionals typically rely on estimating two complex nuisance functions, such as the propensity score and conditional outcome mean for the average treatment effect functional. We consider the problem of how to estimate nuisance functions to obtain optimal rates of convergence for a doubly robust nonparametric functional that has witnessed applications across the causal inference and conditional independence testing literature. For several plug-in estimators and a first-order bias-corrected estimator, we illustrate the interplay between different tuning parameter choices for the nuisance function estimators and sample splitting strategies on the optimal rate of estimating the functional of interest. For each of these estimators and each sample splitting strategy, we show the necessity to either undersmooth or oversmooth the nuisance function estimators under low regularity conditions to obtain optimal rates of convergence for the functional of interest. Unlike the existing literature, we show that plug-in and first-order bias-corrected estimators can achieve minimax rates of convergence across all Hölder smoothness classes of the nuisance functions by careful combinations of sample splitting and nuisance function tuning strategies. We complement these results with numerical simulations illustrating the impact of different nuisance function tuning and sample splitting strategies.

math.ST

On the Asymptotic Inadmissibility of Double Machine Learning Estimators Under Structure-Agnostic Models

Structure-agnostic (SA) models introduced by Balakrishnan et al. (2026) aim to reflect the general lack of knowledge of structural assumptions on data-generating laws such as smoothness or sparsity in practice. Roughly speaking, SA models restrict the observed-data generating law to be in some rn-neighborhood of (black-box machine learning) estimates, treated as given and fixed, where rn encodes the convergence rates of the estimates to the truth. Under SA models, Balakrishnan et al. (2026) show that the popular Double Machine Learning (DML) estimators for three functionals, the quadratic functional in the Gaussian sequence model, the quadratic density integral functional and the expected conditional covariance, are minimax. However, minimax estimators may be inadmissible. In this paper, we show that, for the first two of the three functionals, the DML estimator is asymptotically inadmissible under the SA model. In particular, we show that these two functionals fall into a class of functionals, which we refer to as the monotone bias class. For this class, we exhibit second-order (U-statistic) estimators, which asymptotically dominate DML estimators, under the SA model. These second-order estimators are empirical higher-order influence function (HOIF) estimators introduced in Liu et al. (2017). Furthermore, the empirical HOIF estimator, like the DML estimator, is minimax for the third functional (the expected conditional covariance), although neither asymptotically dominates the other.

math.ST

Pooling Versus Ensembling for Ridge Regression Under Covariate Shift

Datasets in many settings naturally partition into clusters arising from sub-populations, batch effects, or aggregation across multiple sources. A common response to such heterogeneity is to ensemble learners trained on each cluster rather than fit a single model to the pooled data. Prior work motivating such approaches has typically considered settings in which both the covariate distribution and the conditional outcome model differ across clusters; the role of cluster-aware partitioning and ensembling based solely on the covariate distribution remains to be explored. We address this case for ridge-regularized least-squares regression under a linear outcome model and consider all ridge penalty values $λ\geq 0$, including the special case of the ridgeless predictor at $λ= 0$. By considering both fixed-effects and random-effects models, we argue that under random effects, an optimally tuned pooled ridge predictor always outperforms ensembles of individually optimally tuned predictors. For fixed effects, we derive a general formula for the pooled and ensembled predictors to characterize the role of both regression coefficients as well as the predictor distribution shifts. Together, these results generalize prior risk analyses of bagging and random-partition estimation using ridge and ridgeless regression predictors from the i.i.d. setting to encompass covariate shift and heterogeneity-aware partition structure.

math.ST

Toward Joint Prediction of a Longitudinal Marker and a Terminal Event: A bivariate discrete-time framework

Sudden cardiac death (SCD) is a leading cause of death in the U.S. Patients at elevated risk of SCD are primarily treated with an implantable cardioverter-defibrillator (ICD), which may prevent death from cardiovascular causes but may cause severe side effects, such as reduced quality of life from shock-induced pain. Decisions about ICD treatment therefore involve complex personal trade-offs across multiple health events, including mortality and quality of life. While prediction tools could help weigh these trade-offs, they commonly focus on univariate outcomes; at best, they treat other clinical endpoints as inputs, so trade-offs cannot be directly informed. To address this, we propose a novel general Bayesian framework that jointly models a terminal event and a longitudinal marker as a bivariate process over discrete time, for settings where prediction is the primary goal. Discretization of study time lets the framework capture the dynamic interplay between outcomes while avoiding implicit extrapolation beyond truncation by a terminal event. The framework flexibly accommodates phenomena arising in applied contexts, including global time-invariant and local time-dependent dependence structures between the terminal event and the longitudinal marker, and latent association via a shared frailty term. Estimation proceeds via the Bayesian paradigm, yielding patient-specific joint posterior predictions for the time to terminal event and the future marker trajectory. We introduce the framework with a focus on its modeling flexibility, provide guidance on discretization and on Bayesian model construction and selection, and discuss insights from the joint posterior predictions. Finally, we demonstrate the framework's clinical relevance and practicality for SCD and ICD therapy using data from the Sudden Cardiac Death in Heart Failure Trial (SCD-HeFT), an important ICD-related benchmark trial.

stat.ME

Cross-Cluster Weighted Forests

Building trustworthy machine learning algorithms for biological applications requires adapting to data heterogeneity from different sources, batches, distributions, or studies. We propose the 'Cross-Cluster Weighted Forest' (CCWF), an ensembling approach that explicitly leverages heterogeneity in the feature distribution to produce more accurate and more generalizable predictors than the standard Random Forest in cases when data can be naturally clustered. CCWF generalizes the RF architecture to an outer unsupervised layer, supervised subtasks, and ensembling. Specifically it involves unsupervised clustering of the training data, fitting a Random Forest on each cluster, and combining the forests via stacked regression weights that reward cross-cluster generalizability. We provide a theoretical analysis of an analytically tractable forest model showing that cluster-based ensembling is asymptotically more accurate than training a single forest on the full data, with the gain driven by bias reduction. In simulations, we find that CCWF is robust across data-generating regimes and outcome models; furthermore, we explore the influence of data partitioning and ensemble weighting strategies on the benefits of our method. Finally, we apply our approach to cancer molecular profiling and gene expression datasets that are naturally divisible into clusters; in both simulations and real data examples, we illustrate that our approach outperforms classic Random Forest by margins of 30-40%, aligning with our theoretical results. Overall, we show that CCWF provides a statistically grounded prediction algorithm for data spanning multiple domains or sub-populations, a structure common in biological applications.

stat.ML

Sharp minimax risks and phase transitions in sparse submatrix detection

We study the minimax risk for detecting a sparse elevated-mean Gaussian submatrix inside a larger noisy matrix. When the planted submatrix has size $n\times n$ and the ambient matrix has size $N\times N$ with $N = n^{1+α}$, the classical work of \cite{butuceasubmatrix2013} identifies the sharp detection boundary around which the minimax risk converges to $0$ or $1$. This paper extends that zero-one theory by determining the precise asymptotic rate of the minimax risk throughout a two-variable phase diagram. Above the detection boundary, we determine the precise exponent for the stretched or super-exponential decay of the risk. Below the boundary, where the risk tends to 1, we identify the exact polynomial order of the rate of convergence up to absolute multiplicative constants. In both of these regimes, the form of the sharp asymptotics changes around the line $α+ δ= 1/2$ where $δ$ indicates the signed distance from the boundary. Finally, on the detection boundary, we show that the minimax risk converges to the non-degenerate constant $\frac12$ in the very sparse case where $n$ remains fixed and $N \to \infty$. Each of these rates corresponds to the risk of a suitably calibrated scan or sum test, whence follow the upper bounds. To show the sharpness of these bounds, we rely on refined second-moment methods applied to random variables chosen carefully according to the particular regime. Our results also extend to the tensor setting.

math.ST

Robust Causal Inference for EHR-based Studies of Point Exposures with Missingness in Eligibility Criteria

Missingness in variables that define study eligibility criteria is a seldom addressed challenge in electronic health record (EHR)-based settings. It is typically the case that patients with incomplete eligibility information are excluded from analysis without consideration of (implicit) assumptions that are being made, leaving study conclusions subject to potential selection bias. In an effort to ascertain eligibility for more patients, researchers may look back further in time prior to study baseline, and in using outdated values of eligibility-defining covariates may inappropriately be including individuals who, unbeknownst to the researcher, fail to meet eligibility at baseline. To the best of our knowledge, however, very little work has been done to mitigate these concerns. We propose a robust and efficient estimator of the causal average treatment effect on the treated, defined in the study eligible population, in cohort studies where eligibility-defining covariates are missing at random. The approach facilitates the use of flexible machine-learning strategies for component nuisance functions while maintaining appropriate convergence rates for valid asymptotic inference. This method is directly motivated by, and applied throughout to EHR data from Kaiser Permanente to analyze differences between two common bariatric surgical interventions for long-term weight and glycemic outcomes among a cohort of severely obese patients with type II diabetes mellitus.

stat.ME

A Statistical Framework for Understanding Causal Effects that Vary by Treatment Initiation Time in EHR-based Studies

Standard practice in electronic health record (EHR)-based studies evaluating the comparative effectiveness of bariatric surgery relative to no surgery is to estimate and report a constant treatment effect across calendar time. However, real-world treatment strategies can evolve, particularly when comparators include standard of care or surgical procedures where techniques may improve, making it clinically important to ascertain whether efficacy of bariatric surgery has changed over time. Efforts to determine whether treatment efficacy itself is evolving are complicated by changing patient populations, with potential covariate shift in key effect modifiers. Through a comprehensive analysis of EHR data from Kaiser Permanente following two bariatric surgical procedures compared to standard of care, we develop a statistical framework to estimate calendar time-specific average treatment effects and describe both how and why effects vary across treatment initiation time in EHR-based studies. Our approach projects doubly robust, time-specific treatment effect estimates onto candidate marginal structural models and uses a model selection procedure to best describe how effects vary by treatment initiation time. We further introduce a novel summary metric, based on standardization analysis, to quantify the role of covariate shift in explaining observed effect changes and disentangle changes in treatment effects from changes in the patient population receiving treatment.

stat.ME

Semiparametric Efficient Empirical Higher Order Influence Function Estimators

Robins et al. (2008, 2017) applied the theory of higher order influence functions (HOIFs) to derive an estimator of the mean $ψ$ of an outcome Y in a missing data model with Y missing at random conditional on a vector X of continuous covariates; their estimator, in contrast to previous estimators, is semiparametric efficient under the minimal conditions of Robins et al. (2009b), together with an additional (non-minimal) smoothness condition on the density g of X, because the Robins et al. (2008, 2017) estimator depends on a nonparametric estimate of g. In this paper, we introduce a new HOIF estimator that has the same asymptotic properties as the original one, but does not impose any smoothness requirement on g. This is important for two reasons. First, one rarely has the knowledge about the properties of g. Second, even when g is smooth, if the dimension of X is even moderate, accurate nonparametric estimation of its density is not feasible at the sample sizes often encountered in applications. In fact, to the best of our knowledge, this new HOIF estimator remains the only semiparametric efficient estimator of $ψ$ under minimal conditions, despite the rapidly growing literature on causal effect estimation. We also show that our estimator can be generalized to the entire class of functionals considered by Robins et al. (2008) which include the average effect of a treatment on a response Y when a vector X suffices to control confounding and the expected conditional variance of a response Y given a vector X. Simulation experiments are also conducted, which demonstrate that our new estimator outperforms those of Robins et al. (2008, 2017) in finite samples, when g is not very smooth.

math.ST

Asymptotic Inference for Constrained Regression

We consider statistical inference in high-dimensional regression problems under affine constraints on the parameter space. The theoretical study of this is motivated by the study of genetic determinants of diseases, such as diabetes, using external information from mediating protein expression levels. Specifically, we develop rigorous methods for estimating genetic effects on diabetes-related continuous outcomes when these associations are constrained based on external information about genetic determinants of proteins, and genetic relationships between proteins and the outcome of interest. In this regard, we discuss multiple candidate estimators and study their theoretical properties, sharp large sample optimality, and numerical qualities under a high-dimensional proportional asymptotic framework.

stat.ME

Optimal Nuisance Function Tuning for Estimating a Doubly Robust Functional under Proportional Asymptotics

In this paper, we explore the asymptotically optimal tuning parameter choice in ridge regression for estimating nuisance functions of a statistical functional that has recently gained prominence in conditional independence testing and causal inference. Given a sample of size $n$, we study estimators of the Expected Conditional Covariance (ECC) between variables $Y$ and $A$ given a high-dimensional covariate $X \in \mathbb{R}^p$. Under linear regression models for $Y$ and $A$ on $X$ and the proportional asymptotic regime $p/n \to c \in (0, \infty)$, we evaluate three existing ECC estimators and two sample splitting strategies for estimating the required nuisance functions. Since no consistent estimator of the nuisance functions exists in the proportional asymptotic regime without imposing further structure on the problem, we first derive debiased versions of the ECC estimators that utilize the ridge regression nuisance function estimators. We show that our bias correction strategy yields $\sqrt{n}$-consistent estimators of the ECC across different sample splitting strategies and estimator choices. We then derive the asymptotic variances of these debiased estimators to illustrate the nuanced interplay between the sample splitting strategy, estimator choice, and tuning parameters of the nuisance function estimators for optimally estimating the ECC. Our analysis reveals that prediction-optimal tuning parameters (i.e., those that optimally estimate the nuisance functions) may not lead to the lowest asymptotic variance of the ECC estimator -- thereby demonstrating the need to be careful in selecting tuning parameters based on the final goal of inference. Finally, we verify our theoretical results through extensive numerical experiments.

math.ST

Inference on Gaussian mixture models with dependent labels

Gaussian mixture models are widely used to model data generated from multiple latent sources. Despite its popularity, most theoretical research assumes that the labels are either independent and identically distributed, or follows a Markov chain. It remains unclear how the fundamental limits of estimation change under more complex dependence. In this paper, we address this question for the spherical two-component Gaussian mixture model. We first show that for labels with an arbitrary dependence, a naive estimator based on the misspecified likelihood is $\sqrt{n}$-consistent. Additionally, under labels that follow an Ising model, we establish the information theoretic limitations for estimation, and discover an interesting phase transition as dependence becomes stronger. When the dependence is smaller than a threshold, the optimal estimator and its limiting variance exactly matches the independent case, for a wide class of Ising models. On the other hand, under stronger dependence, estimation becomes easier and the naive estimator is no longer optimal. Hence, we propose an alternative estimator based on the variational approximation of the likelihood, and argue its optimality under a specific Ising model.

math.ST

PC Adjusted Testing for Low Dimensional Parameters

In this paper, we investigate the impact of high-dimensional Principal Component (PC) adjustments on inferring the effects of variables on outcomes, with a focus on applications in genetic association studies where PC adjustment is commonly used to account for population stratification. We consider high-dimensional linear regression in the regime where the number of covariates grows proportionally to the number of samples. In this setting, we provide an asymptotically precise understanding of when PC adjustments yield valid tests with controlled Type I error rates. Our results demonstrate that, under both fixed and diverging signal strengths, PC regression often fails to control the Type I error at the desired nominal level. Furthermore, we establish necessary and sufficient conditions for Type I error inflation based on covariate distributions. These theoretical findings are further supported by a series of numerical experiments.

math.ST

Method-of-Moments Inference for GLMs and Doubly Robust Functionals under Proportional Asymptotics

In this paper, we consider the estimation of regression coefficients and signal-to-noise (SNR) ratio in high-dimensional Generalized Linear Models (GLMs), and explore their implications in inferring popular estimands such as average treatment effects in high-dimensional observational studies. Under the ``proportional asymptotic'' regime and Gaussian covariates with known (population) covariance $Σ$, we derive Consistent and Asymptotically Normal (CAN) estimators of our targets of inference through a Method-of-Moments type of estimators that bypasses estimation of high dimensional nuisance functions and hyperparameter tuning altogether. Additionally, under non-Gaussian covariates, we demonstrate universality of our results under certain additional assumptions on the regression coefficients and $Σ$. We also demonstrate that knowing $Σ$ is not essential to our proposed methodology when the sample covariance matrix estimator is invertible. Finally, we complement our theoretical results with numerical experiments and comparisons with existing literature.

math.ST

Sensitivity analysis for nonignorable missing values in blended analysis framework: a study on the effect of bariatric surgery via electronic health records

This paper establishes a series of sensitivity analyses to investigate the impact of missing values in the electronic health records (EHR) that are possibly missing not at random (MNAR). EHRs have gained tremendous interest due to their cost-effectiveness, but their employment for research involves numerous challenges, such as selection bias due to missing data. The blended analysis has been suggested to overcome such challenges, which decomposes the data provenance into a sequence of sub-mechanisms and uses a combination of inverse-probability weighting (IPW) and multiple imputation (MI) under missing at random assumption (MAR). In this paper, we expand the blended analysis under the MNAR assumption and present a sensitivity analysis framework to investigate the effect of MNAR missing values on the analysis results. We illustrate the performance of my proposed framework via numerical studies and conclude with strategies for interpreting the results of sensitivity analyses. In addition, we present an application of our framework to the DURABLE data set, an EHR from a study examining long-term outcomes of patients who underwent bariatric surgery.

stat.ME

Causal Quantile Treatment Effects with missing data by double-sampling

Causal weighted quantile treatment effects (WQTE) are a useful complement to standard causal contrasts that focus on the mean when interest lies at the tails of the counterfactual distribution. To-date, however, methods for estimation and inference regarding causal WQTEs have assumed complete data on all relevant factors. In most practical settings, however, data will be missing or incomplete data, particularly when the data are not collected for research purposes, as is the case for electronic health records and disease registries. Furthermore, such data sources may be particularly susceptible to the outcome data being missing-not-at-random (MNAR). In this paper, we consider the use of double-sampling, through which the otherwise missing data are ascertained on a sub-sample of study units, as a strategy to mitigate bias due to MNAR data in the estimation of causal WQTEs. With the additional data in-hand, we present identifying conditions that do not require assumptions regarding missingness in the original data. We then propose a novel inverse-probability weighted estimator and derive its asymptotic properties, both pointwise at specific quantiles and uniformly across a range of quantiles over some compact subset of (0,1), allowing the propensity score and double-sampling probabilities to be estimated. For practical inference, we develop a bootstrap method that can be used for both pointwise and uniform inference. A simulation study is conducted to examine the finite sample performance of the proposed estimators. The proposed method is illustrated with data from an EHR-based study examining the relative effects of two bariatric surgery procedures on BMI loss at 3 years post-surgery.

stat.ME

Adjusting for Selection Bias Due to Missing Eligibility Criteria in Emulated Target Trials

Target trial emulation (TTE) is a popular framework for observational studies based on electronic health records (EHR). A key component of this framework is determining the patient population eligible for inclusion in both a target trial of interest and its observational emulation. Missingness in variables that define eligibility criteria, however, presents a major challenge towards determining the eligible population when emulating a target trial with an observational study. In practice, patients with incomplete data are almost always excluded from analysis despite the possibility of selection bias, which can arise when subjects with observed eligibility data are fundamentally different than excluded subjects. Despite this, to the best of our knowledge, very little work has been done to mitigate this concern. In this paper, we propose a novel conceptual framework to address selection bias in TTE studies, tailored towards time-to-event endpoints, and describe estimation and inferential procedures via inverse probability weighting (IPW). Under an EHR-based simulation infrastructure, developed to reflect the complexity of EHR data, we characterize common settings under which missing eligibility data poses the threat of selection bias and investigate the ability of the proposed methods to address it. Finally, using EHR databases from Kaiser Permanente, we demonstrate the use of our method to evaluate the effect of bariatric surgery on microvascular outcomes among a cohort of severely obese patients with Type II diabetes mellitus (T2DM).

stat.ME