SearcharxivSearch

arXiv subjects

Samhita Pal

Publications and source records attributed to Samhita Pal.

At least 19 recordsLinked to original sources

Heterogeneous Effects of Continuous Treatments via Conditional Modified Treatment Policies

For continuous treatments such as drug dose or ventilator intensity, a key clinically actionable question is whether a modest, patient-specific adjustment to the current dose would help or harm, rather than whether to treat at all. Standard estimands such as average or conditional dose-response functions require positivity across a wide range of doses, an assumption that routinely fails in observational clinical data where protocols tie dosing to patient characteristics. We develop a framework for characterizing heterogeneity in the effects of small shifts (``nudges'') of a continuous treatment. We study two estimands: the conditional nudge effect, the expected outcome change if an individual at dose $a$ with covariates $\bx$ had their dose shifted by $\delta$, and the conditional modified treatment policy (CMTP) effect, which averages the nudge over the observed dose distribution given covariates. Both are identified under weak, shift-specific local positivity and exchangeability conditions. Recasting the $\delta$-shift as a two-arm comparison through a duplicated-data construction, we develop weighting and A-learning estimators based on squared-error and negative log-likelihood losses, introduce augmented versions that reduce variance without changing the target, and establish asymptotic normality of proposed estimators. Simulations support the theory, and an analysis of mechanical ventilation data from MIMIC-III identifies patient profiles predicted to benefit from a modest reduction in mechanical power.

stat.ME

B-CALM: Bias-Limited Bayesian Borrowing for RCT-Anchored Treatment Effects under Covariate Mismatch

Randomized controlled trials (RCTs) identify treatment effects in the randomized trial population but are often too small for reliable heterogeneity estimation; observational studies (OS) are larger but confounded and measured on only partially overlapping covariates. We develop Bayesian Calibrated ALignment under covariate Mismatch (B-CALM), a Bayesian borrowing method for RCT-defined conditional average treatment effect (CATE) estimation. B-CALM maps source-specific covariates into a shared latent state, jointly models trial and observational outcome surfaces, and uses baseline-bias and comparative-bias functions to represent how the OS departs from the trial estimand. The comparative-bias prior becomes an explicit sensitivity knob: we prove a finite-feature bias-limited information bound showing that observational contrast information about the trial treatment-effect function is capped by the prior precision of this bias function, and derive an effective-sample-size formula showing that the RCT-equivalent information contributed by the OS saturates as OS sample size grows. The theory also combines a PAC-Bayes trial-risk bound with an integral-probability-metric (IPM) alignment and calibration decomposition that separates RCT empirical risk, latent alignment, and residual calibration of the debiased OS surface. In synthetic, semi-synthetic, and pediatric-obesity external-control studies, B-CALM maintains near-nominal average coverage and low negative transfer while pooled and causal-forest baselines can become overconfident under comparative bias.

stat.ME

Hierarchical Projection for Adaptive Knowledge Transfer

Modern data-driven applications increasingly involve learning from multiple heterogeneous sources, where a target dataset is limited but related information is available across domains. Naively combining these sources can degrade performance when relevance varies or spurious signals are present, posing a fundamental challenge for trustworthy cross-domain learning. We propose Projection Transfer Learning (ProjectionTL), a unified framework that integrates hierarchical Bayesian modeling with adaptive projection for selective knowledge transfer. The key idea is to decouple transfer at two levels: first, we construct a source-guided hierarchical prior that aggregates information across sources using data-driven weights, capturing global alignment between each source and the target; second, we refine this borrowing through a posterior-projection step that operates at the feature level, selectively retaining coordinates that exhibit local agreement with the target signal. This two-stage design enables the method to simultaneously perform source selection and feature selection, thereby mitigating negative transfer while preserving interpretability. ProjectionTL provides a principled approach to integrating heterogeneous data across domains, bridging statistical modeling and modern machine learning paradigms for robust and interpretable transfer. Through simulations and real-world biomedical applications, we demonstrate improved accuracy, stability, and interpretability compared to existing methods. Our framework offers a scalable and generalizable strategy for trustworthy cross-domain learning in high-dimensional settings.

cs.LG

Network-aware IV Regression for Causal Node Discovery and Estimation

Estimating causal effects from high-dimensional, structured exposures is a fundamental challenge in modern applications ranging from neuroscience and finance to environmental science. While the literature has addressed high-dimensional instrumental variable (IV) regression, and separately leveraged graph structure in penalized regression, the integration of both, especially for causal support recovery in the presence of latent confounding, remains unexplored. In this work, we propose a novel two-stage regression framework that incorporates instrumental variables and graph-based regularization to uncover sparse causal effects among network-structured exposures. Our method accommodates both valid and partially invalid instruments, and encourages structural similarity among connected predictors through a graph-fused penalty. We establish non-asymptotic guarantees for estimation accuracy and causal variable selection, and demonstrate that our approach yields improved performance over existing methods that ignore network dependencies or invalid IVs. Applied to ADNI brain imaging and genetic data, our method identifies interpretable causal ROIs associated with cognitive outcomes, underscoring the utility of graph-assisted IV regression in neuroscience and beyond.

stat.ME

Omitted-Variable Sensitivity Analysis for Generalizing Randomized Trials

Randomized controlled trials (RCTs) yield internally valid causal effect estimates, but generalizing these results to target populations with different characteristics requires an untestable selection ignorability assumption: conditional on observed covariates, trial participation must be independent of potential outcomes. This assumption fails when unobserved effect modifiers are distributed differently between trial and target populations. We develop a sensitivity analysis framework for trial generalization grounded in omitted variable bias (OVB). Our key theoretical contribution is an exact decomposition showing that external-validity bias equals moderation strength $\times$ moderator imbalance: (i) how strongly an unobserved variable shifts the treatment effect, times (ii) how differently that variable is distributed across populations after covariate adjustment. We introduce scale-free sensitivity parameters based on partial $R^2$ values, enabling closed-form bounds and benchmarking against observed covariates -- practitioners can assess whether conclusions would change if an unobserved moderator were "as strong as" a particular observed variable. Simulations demonstrate that our bounds achieve nominal coverage and remain conservative under model misspecification, while comparisons with alternative sensitivity frameworks highlight the interpretive advantages of the OVB decomposition.

stat.ME

Improving RCT-Based Treatment Effect Estimation Under Covariate Mismatch via Calibrated Alignment

Randomized controlled trials (RCTs) are the gold standard for estimating treatment effects, yet they are often underpowered for detecting effect heterogeneity. Large observational studies (OS) can supplement RCTs for conditional average treatment effect (CATE) estimation, but a key barrier is covariate mismatch: the two sources measure different, only partially overlapping, covariates. We propose CALM (Calibrated ALignment under covariate Mismatch), which learns embeddings that map each source's features into a common representation space. OS outcome models are transferred to the RCT embedding space and calibrated using trial data, preserving causal identification from randomization. Finite-sample risk bounds decompose into alignment error, outcome-model complexity, and calibration complexity terms, making explicit when the learned embedding is accurate enough to reduce variance. We instantiate CALM in two forms: a closed-form linear version, CALM-Lin, and a neural representation-learning version, CALM-NN. Across 51 simulation settings, calibration-based linear methods are effectively tied in linear-CATE regimes, while CALM-NN wins all 22 nonlinear-CATE settings by wide margins. Moreover, on two real-data studies CALM-NN delivers the largest gains over the trial-only baseline.

cs.LG

Improving RCT-Based CATE Estimation Under Covariate Mismatch via Double Calibration

We develop estimators that improve precision of heterogeneous treatment effect estimates that allow borrowing information from observational studies when the available covariates in each data source do not perfectly match. Standard data-borrowing methods often assume perfectly matched covariates. We propose MR-OSCAR, an RCT-calibrated, two-stage estimation approach that first predicts the trial-missing variables using the observational data via imputation and then calibrates observational outcome predictions to the randomized trial, preserving the causal contrast, unlike the results for generalization, where imputation does not improve performance. Our theory gives finite-sample guarantees with a transparent error decomposition including an imputation error that shrinks as the observational mapping becomes more predictable. Simulations show that imputation almost always outperforms naively using only the shared covariates and clarifies when borrowing helps (strong predictability of the missing block, moderate trial size) and when it does not (poor predictability or dominant trial-only moderators). We motivate the approach with the Greenlight Plus trial on early childhood obesity and outline a forthcoming EHR analysis at Vanderbilt, highlighting the use of our method in common scenarios where data do not perfectly align.

stat.ME

Noise-Calibrated Inference from Differentially Private Sufficient Statistics in Exponential Families

Many differentially private (DP) data release systems either output DP synthetic data and leave analysts to perform inference as usual, which can lead to severe miscalibration, or output a DP point estimate without a principled way to do uncertainty quantification. This paper develops a clean and tractable middle ground for exponential families: release only DP sufficient statistics, then perform noise-calibrated likelihood-based inference and optional parametric synthetic data generation as post-processing. Our contributions are: (1) a general recipe for approximate-DP release of clipped sufficient statistics under the Gaussian mechanism; (2) asymptotic normality, explicit variance inflation, and valid Wald-style confidence intervals for the plug-in DP MLE; (3) a noise-aware likelihood correction that is first-order equivalent to the plug-in but supports bootstrap-based intervals; and (4) a matching minimax lower bound showing the privacy distortion rate is unavoidable. The resulting theory yields concrete design rules and a practical pipeline for releasing DP synthetic data with principled uncertainty quantification, validated on three exponential families and real census data.

cs.LG

Sharp Bounds for Treatment Effect Generalization under Outcome Distribution Shift

Generalizing treatment effects from a randomized trial to a target population requires the assumption that potential outcome distributions are invariant across populations after conditioning on observed covariates. This assumption fails when unmeasured effect modifiers are distributed differently between trial participants and the target population. We develop a sensitivity analysis framework that bounds how much conclusions can change when this transportability assumption is violated. Our approach constrains the likelihood ratio between target and trial outcome densities by a scalar parameter $\Lambda \geq 1$, with $\Lambda = 1$ recovering standard transportability. For each $\Lambda$, we derive sharp bounds on the target average treatment effect -- the tightest interval guaranteed to contain the true effect under all data-generating processes compatible with the observed data and the sensitivity model. We show that the optimal likelihood ratios have a simple threshold structure, leading to a closed-form greedy algorithm that requires only sorting trial outcomes and redistributing probability mass. The resulting estimator runs in $O(n \log n)$ time and is consistent under standard regularity conditions. Simulations demonstrate that our bounds achieve nominal coverage when the true outcome shift falls within the specified $\Lambda$, provide substantially tighter intervals than worst-case bounds, and remain informative across a range of realistic violations of transportability.

stat.ME

Causal Discovery with Mixed Latent Confounding via Precision Decomposition

We study causal discovery from observational data in linear Gaussian systems affected by \emph{mixed latent confounding}, where some unobserved factors act broadly across many variables while others influence only small subsets. This setting is common in practice and poses a challenge for existing methods: differentiable and score-based DAG learners can misinterpret global latent effects as causal edges, while latent-variable graphical models recover only undirected structure. We propose \textsc{DCL-DECOR}, a modular, precision-led pipeline that separates these roles. The method first isolates pervasive latent effects by decomposing the observed precision matrix into a structured component and a low-rank component. The structured component corresponds to the conditional distribution after accounting for pervasive confounders and retains only local dependence induced by the causal graph and localized confounding. A correlated-noise DAG learner is then applied to this deconfounded representation to recover directed edges while modeling remaining structured error correlations, followed by a simple reconciliation step to enforce bow-freeness. We provide identifiability results that characterize the recoverable causal target under mixed confounding and show how the overall problem reduces to well-studied subproblems with modular guarantees. Synthetic experiments that vary the strength and dimensionality of pervasive confounding demonstrate consistent improvements in directed edge recovery over applying correlated-noise DAG learning directly to the confounded data.

cs.LG

Efficacy Analysis in Clinical Trials: A Comprehensive Review of Statistical and Machine Learning Approaches

Efficacy testing is a cornerstone of clinical trials, ensuring that medical interventions achieve their intended therapeutic effects. Over the decades, a wide range of statistical methodologies have been developed to address the complexities of clinical trial data, including parametric, nonparametric, Bayesian, and machine learning approaches. Parametric methods, such as t-tests, ANOVA, and LMMs, have traditionally been the foundation of efficacy testing due to their efficiency under well-defined assumptions. Nonparametric techniques, including the Friedman test, Brunner-Munzel test, and modern extensions like nparLD, have emerged as robust alternatives, particularly for skewed, ordinal, or non-normal data. Bayesian methodologies have enabled the incorporation of prior information and uncertainty quantification, while machine learning techniques, such as deep learning and reinforcement learning, are revolutionizing trial designs and outcome predictions. Despite these advancements, significant gaps remain, including challenges in handling high-dimensional data, missingness, and ensuring equitable efficacy testing across diverse populations. This review provides a comprehensive overview of these statistical methods, highlighting their applications, strengths, limitations, and future directions. By bridging traditional statistical frameworks with modern computational techniques, the field can continue to advance toward more reliable and personalized clinical trial methodologies.

stat.OT

DAG DECORation: Continuous Optimization for Structure Learning under Hidden Confounding

We study structure learning for linear Gaussian SEMs in the presence of latent confounding. Existing continuous methods excel when errors are independent, while deconfounding-first pipelines rely on pervasive factor structure or nonlinearity. We propose \textsc{DECOR}, a single likelihood-based and fully differentiable estimator that jointly learns a DAG and a correlated noise model. Our theory gives simple sufficient conditions for global parameter identifiability: if the mixed graph is bow free and the noise covariance has a uniform eigenvalue margin, then the map from $(\B,\OmegaMat)$ to the observational covariance is injective, so both the directed structure and the noise are uniquely determined. The estimator alternates a smooth-acyclic graph update with a convex noise update and can include a light bow complementarity penalty or a post hoc reconciliation step. On synthetic benchmarks that vary confounding density, graph density, latent rank, and dimension with $n<p$, \textsc{DECOR} matches or outperforms strong baselines and is especially robust when confounding is non-pervasive, while remaining competitive under pervasiveness.

cs.LG

Penalized FCI for Causal Structure Learning in a Sparse DAG for Biomarker Discovery in Parkinson's Disease

Parkinson's disease (PD) is a progressive neurodegenerative disorder that lacks reliable early-stage biomarkers for diagnosis, prognosis, and therapeutic monitoring. While cerebrospinal fluid (CSF) biomarkers, such as alpha-synuclein seed amplification assays (alphaSyn-SAA), offer diagnostic potential, their clinical utility is limited by invasiveness and incomplete specificity. Plasma biomarkers provide a minimally invasive alternative, but their mechanistic role in PD remains unclear. A major challenge is distinguishing whether plasma biomarkers causally reflect primary neurodegenerative processes or are downstream consequences of disease progression. To address this, we leverage the Parkinson's Progression Markers Initiative (PPMI) Project 9000, containing 2,924 plasma and CSF biomarkers, to systematically infer causal relationships with disease status. However, only a sparse subset of these biomarkers and their interconnections are actually relevant for the disease. Existing causal discovery algorithms, such as Fast Causal Inference (FCI) and its variants, struggle with the high dimensionality of biomarker datasets under sparsity, limiting their scalability. We propose Penalized Fast Causal Inference (PFCI), a novel approach that incorporates sparsity constraints to efficiently infer causal structures in large-scale biological datasets. By applying PFCI to PPMI data, we aim to identify biomarkers that are causally linked to PD pathology, enabling early diagnosis and patient stratification. Our findings will facilitate biomarker-driven clinical trials and contribute to the development of neuroprotective therapies.

stat.ME

Ensemble Survival Analysis for Preclinical Cognitive Decline Prediction in Alzheimer's Disease Using Longitudinal Biomarkers

Predicting the risk of clinical progression from cognitively normal (CN) status to mild cognitive impairment (MCI) or Alzheimer's disease (AD) is critical for early intervention in Alzheimer's disease (AD). Traditional survival models often fail to capture complex longitudinal biomarker patterns associated with disease progression. We propose an ensemble survival analysis framework integrating multiple survival models to improve early prediction of clinical progression in initially cognitively normal individuals. We analyzed longitudinal biomarker data from the Alzheimer's Disease Neuroimaging Initiative (ADNI) cohort, including 721 participants, limiting analysis to up to three visits (baseline, 6-month follow-up, 12-month follow-up). Of these, 142 (19.7%) experienced clinical progression to MCI or AD. Our approach combined penalized Cox regression (LASSO, Elastic Net) with advanced survival models (Random Survival Forest, DeepSurv, XGBoost). Model predictions were aggregated using ensemble averaging and Bayesian Model Averaging (BMA). Predictive performance was assessed using Harrell's concordance index (C-index) and time-dependent area under the curve (AUC). The ensemble model achieved a peak C-index of 0.907 and an integrated time-dependent AUC of 0.904, outperforming baseline-only models (C-index 0.608). One follow-up visit after baseline significantly improved prediction accuracy (48.1% C-index, 48.2% AUC gains), while adding a second follow-up provided only marginal gains (2.1% C-index, 2.7% AUC). Our ensemble survival framework effectively integrates diverse survival models and aggregation techniques to enhance early prediction of preclinical AD progression. These findings highlight the importance of leveraging longitudinal biomarker data, particularly one follow-up visit, for accurate risk stratification and personalized intervention strategies.

stat.AP

Bayesian High-dimensional Linear Regression with Sparse Projection-posterior

We consider a novel Bayesian approach to estimation, uncertainty quantification, and variable selection for a high-dimensional linear regression model under sparsity. The number of predictors can be nearly exponentially large relative to the sample size. We put a conjugate normal prior initially disregarding sparsity, but for making an inference, instead of the original multivariate normal posterior, we use the posterior distribution induced by a map transforming the vector of regression coefficients to a sparse vector obtained by minimizing the sum of squares of deviations plus a suitably scaled $\ell_1$-penalty on the vector. We show that the resulting sparse projection-posterior distribution contracts around the true value of the parameter at the optimal rate adapted to the sparsity of the vector. We show that the true sparsity structure gets a large sparse projection-posterior probability. We further show that an appropriately recentred credible ball has the correct asymptotic frequentist coverage. Finally, we describe how the computational burden can be distributed to many machines, each dealing with only a small fraction of the whole dataset. We conduct a comprehensive simulation study under a variety of settings and found that the proposed method performs well for finite sample sizes. We also apply the method to several real datasets, including the ADNI data, and compare its performance with the state-of-the-art methods. We implemented the method in the \texttt{R} package called \texttt{sparseProj}, and all computations have been carried out using this package.

stat.ME

Bayesian High-dimensional Grouped-regression using Sparse Projection-posterior

We present a novel Bayesian approach for high-dimensional grouped regression under sparsity. We leverage a sparse projection method that uses a sparsity-inducing map to derive an induced posterior on a lower-dimensional parameter space. Our method introduces three distinct projection maps based on popular penalty functions: the Group LASSO Projection Posterior, Group SCAD Projection Posterior, and Adaptive Group LASSO Projection Posterior. Each projection map is constructed to immerse dense posterior samples into a structured, sparse space, allowing for effective group selection and estimation in high-dimensional settings. We derive optimal posterior contraction rates for estimation and prediction, proving that the methods are model selection consistent. Additionally, we propose a Debiased Group LASSO Projection Map, which ensures exact coverage of credible sets. Our methodology is particularly suited for applications in nonparametric additive models, where we apply it with B-spline expansions to capture complex relationships between covariates and response. Extensive simulations validate our theoretical findings, demonstrating the robustness of our approach across different settings. Finally, we illustrate the practical utility of our method with an application to brain MRI volume data from the Alzheimer's Disease Neuroimaging Initiative (ADNI), where our model identifies key brain regions associated with Alzheimer's progression.

stat.ME

Discovering Candidate Genes Regulated by GWAS Signals in Cis and Trans

Understanding the genetic underpinnings of complex traits and diseases has been greatly advanced by genome-wide association studies (GWAS). However, a significant portion of trait heritability remains unexplained, known as ``missing heritability". Most GWAS loci reside in non-coding regions, posing challenges in understanding their functional impact. Integrating GWAS with functional genomic data, such as expression quantitative trait loci (eQTLs), can bridge this gap. This study introduces a novel approach to discover candidate genes regulated by GWAS signals in both cis and trans. Unlike existing eQTL studies that focus solely on cis-eQTLs or consider cis- and trans-QTLs separately, we utilize adaptive statistical metrics that can reflect both the strong, sparse effects of cis-eQTLs and the weak, dense effects of trans-eQTLs. Consequently, candidate genes regulated by the joint effects can be prioritized. We demonstrate the efficiency of our method through theoretical and numerical analyses and apply it to adipose eQTL data from the METabolic Syndrome in Men (METSIM) study, uncovering genes playing important roles in the regulatory networks influencing cardiometabolic traits. Our findings offer new insights into the genetic regulation of complex traits and present a practical framework for identifying key regulatory genes based on joint eQTL effects.

q-bio.GN

Coverage of Credible Sets for Regression under Variable Selection

We study the asymptotic frequentist coverage of credible sets based on a novel Bayesian approach for a multiple linear regression model under variable selection. We initially ignore the issue of variable selection, which allows us to put a conjugate normal prior on the coefficient vector. The variable selection step is incorporated directly in the posterior through a sparsity-inducing map and uses the induced prior for making an inference instead of the natural conjugate posterior. The sparsity-inducing map minimizes the sum of the squared l2-distance weighted by the data matrix and a suitably scaled l1-penalty term. We obtain the limiting coverage of various credible regions and demonstrate that a modified credible interval for a component has the exact asymptotic frequentist coverage if the corresponding predictor is asymptotically uncorrelated with other predictors. Through extensive simulation, we provide a guideline for choosing the penalty parameter as a function of the credibility level appropriate for the corresponding coverage. We also show finite-sample numerical results that support the conclusions from the asymptotic theory. We also provide the credInt package that implements the method in R to obtain the credible intervals along with the posterior samples.

stat.ME