SearcharxivSearch

arXiv subjects

Anik Burman

Publications and source records attributed to Anik Burman.

5 recordsLinked to original sources

Transporting Randomized Trial Effects to Real-World Populations via Riesz-Calibrated Optimal Transport

Randomized trials support causal inference, but differences between trial and target populations can limit the transportability of treatment effects to real-world settings. Many existing approaches model the propensity of trial participation and can therefore be sensitive to model misspecification and weak overlap of the covariate distributions. Optimal Transport (OT) offers a different route by comparing the trial and target populations directly in covariate space. We develop RICOT, a Riesz-calibrated OT procedure transporting treatment effects to a treated target population. We consider a semi-unbalanced OT with entropic regularization where the source marginals are relaxed. We show that the uncalibrated OT introduces a bias which does not shrink with increasing sample size. RICOT removes this bias by imposing calibration equations directly within the transport problem. With a growing calibration sieve, the calibrated weight consistently estimates the target-to-trial density ratio, equivalently the Riesz representer of the target expectation functional, even when the entropic and source-relaxation parameters remain fixed and positive. Combined with outcome regression, the resulting estimator is doubly robust and attains the semiparametric efficiency bound under suitable rate conditions. Its variance is estimated directly from the influence function, without resampling or repeated OT optimization. Simulations show low bias and near-nominal coverage across a range of overlap and misspecification settings, including settings in which sampling-score methods perform poorly. We illustrate RICOT in a real-world application involving a rare progressive cardiomyopathy, comparing conventional IPW and AIPW estimators with our OT-based IPW and doubly robust estimators for transporting the randomized treatment effect to a real-world population receiving the same treatment.

stat.ME

Doubly-Unlinked Regression for Dependent Data

Shuffled regression concerns settings in which covariates and responses are observed without their correct pairing. In dependent-data problems, a second form of missing correspondence can arise when responses are also detached from the latent temporal, spatial, or geometric domain that induces their dependence structure. We study regression under this joint loss of correspondence and, to our knowledge, provide the first systematic treatment of this setting. Specifically, we consider a doubly-unlinked regression model in which both the covariate-response link and the response-domain link are unknown, represented by two latent permutation matrices, while dependence is induced by an unobserved stochastic process. This framework unifies shuffled regression and latent-domain permutation models within a common dependent-data setting. We characterize signal-to-noise regimes governing recovery of the regression parameter and the latent permutations, and show that consistent estimation of the regression coefficient can be achieved under strictly weaker conditions than exact permutation recovery. To address the combinatorial difficulty of inference, we develop REPAIR, a variational Bayes method based on a block-structured permutation model that captures localized scrambling while substantially reducing computational complexity. Simulations and an applied example illustrate the empirical behavior of REPAIR and support the theoretical results.

math.ST

Robust Spatial Confounding Adjustment via Basis Voting

Estimating effects of spatially structured exposures is complicated by unmeasured spatial confounders, which undermine identifiability in spatial linear regression models unless structural assumptions are imposed. We develop a general framework for effect estimation in spatial regression models that relaxes the commonly assumed requirement that exposures contain higher-frequency variation than confounders. We propose basis voting, a plurality-rule estimator - novel in the spatial literature - that consistently identifies causal effects only under the assumption that, in a spatial basis expansion of the exposure and confounder, there exist several basis functions in the support of the exposure but not the confounder. This assumption generalizes existing assumptions of differential basis support used for identification of the causal effect under spatial confounding, and does not require prior knowledge of which basis functions satisfy this support condition. We design this estimator as the mode of several candidate estimators each computed based on a single working basis function. We also show that the standard projection-based candidate estimator typically used in other plurality-rule based methods is inefficient, and provide a more efficient novel candidate. Extensive simulations and a real-world application demonstrate that our approach reliably recovers unbiased causal estimates whenever exposure and confounder signals are separable on a plurality of basis functions. By not relying on higher-frequency variation, our method remains applicable to settings where exposures are smooth spatial functions, such as distance to pollution sources or major roadways, common in environmental studies.

stat.ME

High-dimensional Portfolio Optimization using Joint Shrinkage

We consider the problem of optimizing a portfolio of financial assets, where the number of assets can be much larger than the number of observations. The optimal portfolio weights require estimating the inverse covariance matrix of excess asset returns, classical solutions of which behave badly in high-dimensional scenarios. We propose to use a regression-based joint shrinkage method for estimating the partial correlation among the assets. Extensive simulation studies illustrate the superior performance of the proposed method with respect to variance, weight, and risk estimation errors compared with competing methods for both the global minimum variance portfolios and Markowitz mean-variance portfolios. We also demonstrate the excellent empirical performances of our method on daily and monthly returns of the components of the S&P 500 index.

q-fin.PM

An AI-Enabled Agent-Based Simulation Platform for Studying COVID-19 Pandemic

Understanding outbreak dynamics is essential for designing effective control measures. We developed an agent-based model to examine how changes in epidemiological and intervention parameters affect infection progression in a synthetic population. The model incorporates individual demographic characteristics, including age, sex, and working status, as well as the number and location of infection epicentres, diagnostic sensitivity, the proportion of asymptomatic infections, and the timing and duration of lockdowns. By tracking each individual, the simulator characterizes infection progression through a community over time. In a closed population of 10000 people, cases peaked around the sixth week and declined by approximately the fifteenth week in the absence of lockdown. When primary cases were introduced within densely populated clusters, cases peaked earlier and declined more slowly. Lockdowns delayed and reduced the infection peak, whereas lower diagnostic sensitivity increased cases and deaths. The number of cases decreased as the proportion of asymptomatic infections increased under the model's assumptions. The model produces reproducible estimates under realistic parameter settings and can accommodate factors such as infectivity period, testing yield, socioeconomic status, daily travel, awareness, population density, and social distancing. It can also be adapted to infections with similar transmission dynamics. The model is available as an open, interactive web application that enables users without programming experience to design scenarios and examine outbreak dynamics in real time. Beyond forecasting, the simulator provides a reusable in-silico environment, or digital twin, for synthetic-data generation and AI-assisted optimization of intervention policies.

q-bio.PE