Searcharxiv⌕ Search

arXiv subjects

M. Ehsan Karim

Publications and source records attributed to M. Ehsan Karim.

10 recordsLinked to original sources

When bad adjustment looks good: what goes wrong in Plasmode 0.1.0 simulations

Comparisons of confounding-control methods are informative only when simulated data contain the treatment-covariate relationship those methods address. We assessed whether Plasmode 0.1.0 preserves that relationship, uses the returned exposure to generate outcomes, and reports the effect implied by its generating model. We inspected the source and ran the generator under identical seeds with source rows as supplied, sorted by observed exposure, or using an alignment correction. We ran 200 null-effect replicates per strategy in a synthetic cohort and the SUPPORT/Right Heart Catheterisation cohort. A sampled subject usually received another subject's exposure probability. In the synthetic cohort, source-model AUC was 0.503 before and 0.841 after alignment. Crude risk-difference bias moved from +0.0051 to +0.3037. Other orders preserved or reversed the association. After alignment, deliberately misspecified estimators retained risk-difference biases of 0.10-0.11. Correctly specified estimators were nearly unbiased. With the unmodified generator, all five appeared similarly unbiased. At an injected odds ratio of 2, the conditional log odds ratio appeared on the observed exposure (+0.696; target 0.693), not the returned exposure (+0.017). Misalignment can weaken, preserve, or reverse designed confounding and thereby distort method comparisons. Alignment correction does not repair outcome-exposure decoupling, disagreement between the reported and generating effects, or marginal miscalibration.

stat.ME↗

A Plasmode Benchmark Can Refute but Cannot Certify: Projection and Regularity in Causal-Inference Simulation

Plasmode simulation - resampling covariates, and often treatment, from a real cohort and regenerating the outcome from a model with an injected effect - is a standard tool for comparing causal-inference estimators. We argue that its validity turns on two separable design choices, and that separating them clarifies a recent debate over resampling versus regenerating the treatment. Our central message limits what such a benchmark can show: a plasmode can refute a universal claim about an estimator but cannot certify one. Refuting needs only that the simulated data genuinely belong to the class over which a guarantee is quoted; certifying also needs the benchmark's difficulty to carry over to the target population, which the analyst cannot verify. Two choices set that difficulty. Regularity is whether the simulated law preserves the overlap that inverse-weighting and targeted-learning theory require. Resampling the real treatment with fine-grained covariates destroys it, and the resulting coverage failure falls on flexible, machine-learning-based estimators, while parametric-nuisance estimators - including inverse-probability weighting - stay near nominal; it disappears once the treatment is regenerated from a bounded propensity or the covariates are coarsened, and cross-fitting removes only part of it. Projection is that the injected outcome model hands a matching estimator the right answer by construction; this survives the regularity fix and can reverse estimator rankings even when overlap is intact. We give a regular-plasmode algorithm, a confirming real-data simulation, and practical recommendations. A plasmode is a filter against estimators, not a warrant for them.

stat.ME↗

Three Routes to One Answer: Reconciling AIPW, TMLE, and Double Machine Learning for Applied Researchers

Augmented inverse-probability weighting (AIPW), targeted maximum likelihood estimation (TMLE), and double/debiased machine learning (DML) are three routes to the same efficient influence function for the average treatment effect --- settled theory we treat as background. This tutorial's contribution is its worked, shared-nuisance reconciliation on real data: what a practitioner must actually match for the routes, and the software packages, to agree. Working the effect of smoking cessation on weight change in the open NHEFS data (n=1566) with one shared Super Learner library and identical cross-fitting folds, we build all three estimators by hand from one influence function, in full-sample, cross-fit, and double-cross-fit variants; the six resulting doubly-robust estimates span only 3.32--3.42 kg, consistent with the established benchmark. We then reconcile the same estimand across our engine and the tmle, AIPW, DoubleML, and tmle3 packages. At their defaults the estimates span 3.32--3.49 kg. Once the library and folds are matched and single-split noise is averaged out, the three library-sharing implementations agree to within 0.01 kg --- so the residual spread traces to the nuisance library, folds, and repetitions, not to the estimator label. Because the estimators share one influence function, they agree under good overlap; under a positivity violation the pooled ATE is not identified without additional extrapolation assumptions, and their finite-sample estimates can then diverge sharply. We therefore place a positivity diagnosis ahead of estimator choice, illustrate the failure on a no-overlap example, and close with a reporting checklist. Open-source R code reproduces every number.

stat.ME↗

Best for which estimand? A known-truth benchmark of longitudinal-matching and target-trial-emulation methods for time-varying treatments

On a non-collapsible survival mechanism, longitudinal-matching and target-trial-emulation methods are not competing estimators of one truth but answers to different causal questions, so a benchmark that scores them against a single "true hazard ratio" fabricates bias. We provide the direct comparison of relative efficiency, variance estimation, and model sensitivity that reviews find lacking. On a deliberately non-collapsible continuous-time Cox mechanism with known truth, the dominant families (sequential Cox, sequential stratification, risk-set matching, and inverse-probability-of-treatment-weighted (IPTW) marginal structural models) target numerically distinct causal estimands (marginal, conditional, two average-treatment-effect-on-the-treated, and intention-to-treat versus per-protocol). First, we quantify the phantom bias a shared marginal truth fabricates: 0.32-0.33 log-cumulative-hazard-ratio units for the matching estimators and 0.15 for the conditional method; the associational naive time-dependent Cox sits 0.76 away, a total discrepancy compounding the estimand gap with confounding. Second, a rank reversal: the recommended method flips with the target estimand, and a low-variance off-target estimator can still win on mean-squared error. Third, a cross-family variance result: the cluster-robust sandwich is closer to nominal for the trial-stacking estimator (0.90) but under-covers the matching estimators (0.77-0.82), which a prespecified n=500 bootstrap sub-study brings to 0.95-0.96. Fourth, model sensitivity: omitting a confounder induces 0.45-0.50 log-hazard-ratio bias and undercoverage, and intention-to-treat and per-protocol effects diverge as switching increases; a heart-transplant analysis illustrates these. On a second mechanism three of four findings replicate, the rank reversal attenuating and model sensitivity proving calibration-dependent.

stat.ME↗

Cluster on the Subject, Not the Record: Confidence Intervals and Simultaneous Bands for Additive-Hazards Sequential Trial Emulation

Sequential trial emulation (STE) estimates the effect of a sustained treatment by stacking nested emulated trials with inverse-probability weighting. Additive-hazards STE estimators of the marginal risk difference recommend the nonparametric bootstrap without evaluating its coverage. Using a correctly-specifiable mechanism (up to a small, disclosed residual), we compare analytic and bootstrap standard errors for both estimands. Exploiting the closed-form linearity of the additive-hazards estimating equation, we derive the influence functions and prove the default row-level robust variance inconsistent for the marginal risk-difference curve, and show the same failure empirically for the constant hazard difference: the row-level variance omits a within-subject cross-trial covariance that is positive under a sign condition we verify across our mechanisms, and is anticonservative at every horizon except the first - only there, where the covariance is zero, is the row-level standard error unimpaired. The subject-clustered variance and multiplier bootstrap are consistent for the fixed-weight linearisation and support simultaneous confidence bands, whose measured coverage is 0.88. Across our simulations the model-based and row-level robust intervals are anticonservative and worsen with sample size, coverage falling to 0.71 at n=5000; clustering leaves the constant-hazard-difference coverage near 0.86 at n=5000, and the multiplier bootstrap leaves the risk-difference-curve coverage near 0.90 (0.86 at the longest horizon). The STE constant hazard difference is a design-weighted summary of a time-varying effect, dependent on the trial structure. We illustrate on the Stanford heart transplant data and provide them in the steCI R package.

stat.ME↗

When Does Trial-Real-World Data Fusion Improve Precision? Model Auditing and Selection-Aware Inference for Adaptive-TMLE

Augmenting a randomized controlled trial (RCT) with real-world data (RWD) promises greater efficiency, but how much a given fusion delivers, and how to attach honest uncertainty to that gain, are rarely characterized. Using adaptive targeted maximum likelihood estimation (A-TMLE) as a worked example of an estimator that learns a working model and then debiases it, we develop three reproducible tools for reliable evidence from combined trial and real-world data. First, a report card that makes the data-adaptively learned bias model auditable: on simulated data it measures how well the model recovers the true enrollment-effect surface and attributes the estimator's variance to its structural parts. Second, a map of when fusion helps versus hurts, benchmarked against a matched trial-only estimator; the efficiency gain is driven mainly by the magnitude of the real-world bias rather than its functional complexity (a dominance an exact population-oracle variance identity explains), it crosses break-even near a moderate bias and erodes as the trial grows, so the advantage is finite-sample, not super-efficiency. Third, selection-aware inference for the gain, treated as a data-adaptive estimand: the naive standard error undercovers, and among ten candidate standard errors only a block jackknife achieved consistently near- or above-nominal coverage, though conservatively. Across six fusions of three openly available trials (a biomedical HIV trial, a public-health trial, and a job-training trial), only one interval clears one, and only marginally; in the rest, fusion has not earned an efficiency claim over the RCT alone. On real data the toolkit therefore functions mainly as a guardrail: the learned-model dimension is a stress diagnostic, not a proxy for ground truth, and the block-jackknife interval decides whether fusion or the RCT-only analysis should be primary.

stat.ME↗

Does a Developed Comorbidity Index Really Add Value? A Selection-Aware Bootstrap for Post-Selection Concordance

Disease-specific comorbidity indices are routinely developed by building several candidate constructions and reporting the best-scoring one, then claiming it adds discriminative value over a fixed off-the-shelf comparator such as the Charlson or Elixhauser score. We show that the optimism correction in standard use does not make that claim valid. Because it corrects the selected model as if it were the only one ever fit, it omits the winner's-curse term from choosing the best of several candidates; so its confidence interval for the incremental concordance is not merely optimistic in small samples but structurally miscalibrated, and does not shrink as the sample grows. At a true null it inflates false claims of added value above the nominal level, increasingly so as more candidates are screened. We introduce a drop-in selection-aware bootstrap that re-runs the best-of-several selection inside each resample with the comparator held fixed, removing the structural bias. In a fully-known-truth simulation, 95% coverage under the standard correction falls from 0.94 with one candidate to 0.70 with a hundred, while the selection-aware interval holds near nominal; its coverage matches a calibrated cross-validation interval, and at a matched error rate it is at least as powerful. The results hold under Uno's concordance, and a semi-synthetic experiment on real survey data confirms when the correction matters. In practice, if several constructions were tried, report a selection-aware interval, most needed with many similar-quality candidates and few events per candidate. The scope is discrimination only; software and results reproduce every finding.

stat.ME↗

When Does Survey-Aware Cross-Validation Matter? The ICC, Not the Design Effect

K-fold cross-validation assumes exchangeable observations, violated by the stratification, clustering, and unequal weights of complex sample surveys. Design-respecting "survey CV" exists, but the question of when the extra care changes any conclusion has remained open. We answer it before validating any model, with two inexpensive diagnostics: the within-cluster intraclass correlation (ICC) of the outcome and of a preliminary linear predictor. A simulated positive control demonstrates their sensitivity, with naive cross-validation growing optimistic about new-cluster performance as the ICC rises while cluster-level folds stay honest. We then evaluate paired naive-versus-design-respecting cross-validation in three national health surveys (chronic-pain, diabetes, and adolescent-suicidality prediction; penalized and random-forest learners, plus an unpenalized comparator) - one reanalysis, one prospective application, and one prespecified screen. In all three the diagnostics correctly anticipated the outcome: no scheme difference of practical size, and the only interval excluding zero showed pessimism, not the optimism that cluster leakage produces - even where the design effect was large (which, unlike the ICC, is not the right trigger). We also show the stratified recipe is often infeasible in public-use designs and give a fallback hierarchy, and we document weight-handling errors whose order-of-magnitude artifacts dwarfed any fold-scheme effect. Reproducible code accompanies the paper.

stat.ME↗

Which Regularized Propensity-Score and Doubly Robust Methods Are Best Calibrated When Exposures or Outcomes Are Rare? A Plasmode Study of Proxy-Based Confounding Adjustment

Purpose. Confounding adjustment in health-care database studies screens large proxy libraries where events per variable are low, straining standard propensity score (PS) methods, and many regularized variable-selection strategies exist (outcome-adaptive LASSO [OAL], group LASSO/GLiDeR, highly adaptive LASSO [HAL]). Yet few comparisons have varied exposure prevalence within such a selection menu, and none pairs it with doubly robust estimation, compute accounting, and a null-RD truth anchor. Methods. We conducted a plasmode simulation anchored on National Health and Nutrition Examination Survey data (2013-2018; 25 investigator-specified covariates, 142 prescription-derived proxies), comparing ten pipelines combining these strategies with inverse probability of treatment weighting (IPTW) and targeted maximum likelihood estimation (TMLE). Three scenarios were evaluated under a known null (true risk difference, RD = 0): frequent, rare-exposure, and rare-outcome. We report bias, standard error (SE), relative error, 95% coverage, and runtime. Results. HAL (G-Computation) had near-zero bias but highly concentrated estimates, giving near-unity coverage and large relative error (106-186%). OAL (IPTW), GLiDeR, and HAL (TMLE) were best calibrated, whereas the regularized-LASSO TMLE pipelines under-covered modestly (91-93%) in the rare scenarios. Under rare exposure, LASSO-IPTW had the largest bias and inflated SE and over-covered (conservatively), problems that TMLE removed. On real data, methods agreed (RD approximately 0.07-0.085). Runtimes spanned <1 s to >16 h. Conclusions. Under a null benchmark, pairing outcome-aware selection (OAL, GLiDeR) or doubly robust estimation (TMLE) with regularized models best balanced bias, calibration, and robustness to rarity. The rare-exposure arm exposed the largest gaps; method choice should weigh the prioritized metric against compute.

stat.AP↗

Cross-Fitted Survey-Weighted TMLE with Design-Based Variance for Causal Machine Learning

Cross-fitting is not a refinement of survey-weighted causal machine learning but, once the nuisances are flexible, what restores valid inference. We study the population average treatment effect under a stratified multistage design, estimated by a survey-aware targeted maximum likelihood estimator (TMLE) whose variance is obtained by Taylor-series linearization of the influence function, treating the primary sampling unit as the replication unit. Our central result is that this validity turns on cross-fitting at the cluster level: sufficiency is established in theory, and the failure without it is shown in simulation. Once flexible learners cross a complexity (Donsker) boundary, single-fit survey TMLE can severely under-cover, and internal cluster-aware cross-validation does not substitute for cross-fitting; among the estimators we evaluate, only out-of-fold fitting at the cluster level restores valid coverage. In simulations spanning a many-PSU and an NHANES-like design, on a diverse ensemble the single-fit and internal cross-validation estimators cover at about 0.89-0.91 and 0.85-0.88 while the cross-fitted estimator holds at 0.93-0.95, and an aggressively grown learner drives single-fit coverage to 0.18-0.22. Two scope choices are deliberate: survey-weighted point estimation is prior work, and the nuisance product-rate condition is assumed and probed empirically. Within these conditions we prove asymptotic normality and design-consistency of the linearization variance. Four NHANES analyses and open-source software illustrate the method.

stat.ME↗