Searcharxiv⌕ Search

arXiv subjects

Shomoita Alam

Publications and source records attributed to Shomoita Alam.

6 recordsLinked to original sources

Framing Causal Questions in Sports Analytics: A Tutorial on Estimand Choice Illustrated Through Crossing in Soccer

Causal inference has become an accepted analytic framework in sports analytics, where experimentation is rarely feasible. A key consideration is the choice of estimand, specifically, whether to target the Average Treatment Effect (ATE), which reflects the effect of an action across the entire population, or the Average Treatment Effect on the Treated (ATT), which reflects the effect among those who actually took the action. Using data from nearly all 240 matches of the 2019 Chinese Super League season, we apply propensity score matching to estimate the causal effect of crossing on shot creation in soccer. The ATE and ATT are nearly identical (0.033 and 0.035 respectively), a result we attribute to substantial overlap in propensity score distributions between plays where a cross was and was not attempted. To illustrate when these estimands diverge, we construct two simulation scenarios with known ground truth: one reproducing the high-overlap structure of the real data, where ATE and ATT coincide, and one engineered to exhibit severe confounding and low overlap, where they diverge substantially. While empirical findings are specific to the 2019 Chinese Super League season, the case study and simulations provide a principled guide to estimand choice in causal analyses of sports data.

stat.AP↗

Anchoring Convenience Survey Samples to a Baseline Census for Vaccine Coverage Monitoring in Global Health

While conducting probabilistic surveys is the gold standard for assessing vaccine coverage, implementing these surveys poses challenges for global health. There is a need for more convenient option that is more affordable and practical. Motivated by childhood vaccine monitoring programs in rural areas of Chad and Niger, we conducted a simulation study to evaluate calibration-weighted design-based and logistic regression-based imputation estimators of the finite-population proportion of MCV1 coverage. These estimators use a hybrid approach that anchors non-probabilistic follow-up survey to probabilistic baseline census to account for selection bias. We explored varying degrees of non-ignorable selection bias (odds ratios from 1.0-1.5), percentage of villages sampled (25-75%), and village-level survey response rate to the follow-up survey (50-80%). Our performance metrics included bias, coverage, and proportion of simulated 95% confidence intervals falling within equivalence margins of 5% and 7.5% (equivalence tolerance). For both adjustment methods, the performance worsened with higher selection bias and lower response rate and generally improved as a larger proportion of villages was sampled. Under the worst scenario with 1.5 OR, 25% village sampled, and 50% survey response rate, both methods showed empirical biases of 2% or less, below 95% coverage, and low equivalence tolerances. In more realistic scenarios, the performance of our estimators showed lower biases and close to 95% coverage. For example, at OR$\leq$1.2, both methods showed high performance, except at the lowest village sampling and participation rates. Our simulations show that a hybrid anchoring survey approach is a feasible survey option for vaccine monitoring.

stat.AP↗

Simulation-Guided Planning of a Target Trial Emulated Cluster Randomized Trial for Mass Small-Quantity Lipid Nutrient Supplementation Combined with Expanded Program on Immunization in Rural Niger

Background: Target trial emulation (TTE) that applies trial design principles to improve the analysis of non-randomized studies is increasingly being used. Applications of TTE to emulate cluster randomized trials (RCTs) have been limited. This study explored how to integrate simulation-guided design into the TTE framework to inform planning of a non-randomized cluster trial. Methods: We performed simulations to prospectively plan data collection of a non-randomized study emulating a village-level cluster RCT when cluster-randomization was infeasible. The planned study will assess the impact of mass distribution of nutritional supplements embedded within an existing immunization program to improve pentavalent vaccination rates among children 12-24 months old in Niger. The design included covariate-constrained random selection of villages for outcome ascertainment at follow-up. Simulations used baseline census data on pentavalent vaccination rates and cluster-level covariates to compare the type I error rate and power of four statistical methods: beta-regression; quasi-binomial regression; inverse probability of treatment weighting (IPTW); and naive Wald test. Results: Of the four analytic methods considered, only IPTW and beta-regression controlled the type I error rate at 0.05, but IPTW yielded poor statistical power. Beta-regression that showed adequate statistical power was chosen as our primary analysis. Conclusions: Adopting simulation-guided design principles within TTE can enable robust planning of a group-level non-randomized study emulating a cluster RCT. Lessons from this study also apply to TTE planning of individually-RCTs.

stat.AP↗

Comparison of Simulation-Guided Design to Closed-Form Power Calculations in Planning a Cluster Randomized Trial with Covariate-Constrained Randomization: A Case Study in Rural Chad

Current practices for designing cluster-randomized trials (cRCTs) typically rely on closed-form formulas for power calculations. For cRCTs using covariate-constrained randomization, the utility of conventional calculations might be limited, particularly when data is nested. We compared simulation-based planning of a nested cRCT using covariate-constrained randomization to conventional power calculations using OptiMAx-Chad as a case study. OptiMAx-Chad will examine the impact of embedding mass distribution of small-quantity lipid-based nutrient supplements within an expanded programme on immunization on first-dose measles-containing vaccine (MCV1) coverage among children aged 12-24 months in rural villages in Ngouri. Within the 12 health areas to be randomized, a random subset of villages will be selected for outcome collection. 1,000,000 assignments of health areas with different possible village selections were generated using covariate-constrained randomization to balance baseline village characteristics. The empirically estimated intracluster correlation coefficient (ICC) and the World Health Organization (WHO) recommended values of 1/3 and 1/6 were considered. The desired operating characteristics were 80% power at 0.05 one-sided type I error rate. Using conventional calculations target power for a realistic treatment effect could not be achieved with the WHO recommended values. Conventional calculations also showed a plateau in power after a certain cluster size. Our simulations matched the design of OptiMAx-Chad with covariate adjustment and random selection, and showed that power did not plateau. Instead, power increased with increasing cluster size. Planning complex cRCTs with covariate constrained randomization and a multi-nested data structure with conventional closed-form formulas can be misleading. Simulations can improve the planning of cRCTs.

stat.AP↗

Estimands and Their Implications for Evidence Synthesis for Oncology: A Simulation Study of Treatment Switching in Meta-Analysis

The ICH E9(R1) addendum provides guidelines on accounting for intercurrent events in clinical trials using the estimands framework. However, there has been limited attention to the estimands framework for meta-analysis. Using treatment switching, a well-known intercurrent event that occurs frequently in oncology, we conducted a simulation study to explore the bias introduced by pooling together estimates targeting different estimands in a meta-analysis of randomized clinical trials (RCTs) that allowed treatment switching. We simulated overall survival data of a collection of RCTs that allowed patients in the control group to switch to the intervention treatment after disease progression under fixed-effects and random-effects models. For each RCT, we calculated effect estimates for a treatment policy estimand that ignored treatment switching, and a hypothetical estimand that accounted for treatment switching either by fitting rank-preserving structural failure time models or by censoring switchers. Then, we performed random-effects and fixed-effects meta-analyses to pool together RCT effect estimates while varying the proportions of trials providing treatment policy and hypothetical effect estimates. We compared the results of meta-analyses that pooled different types of effect estimates with those that pooled only treatment policy or hypothetical estimates. We found that pooling estimates targeting different estimands results in pooled estimators that do not target any estimand of interest, and that pooling estimates of varying estimands can generate misleading results, even under a random-effects model. Adopting the estimands framework for meta-analysis may improve alignment between meta-analytic results and the clinical research question of interest.

stat.ME↗

Multivariate regression with missing response data for modelling regional DNA methylation QTLs

Identifying genetic regulators of DNA methylation (mQTLs) with multivariate models enhances statistical power, but is challenged by missing data from bisulfite sequencing. Standard imputation-based methods can introduce bias, limiting reliable inference. We propose \texttt{missoNet}, a novel convex estimation framework that jointly estimates regression coefficients and the precision matrix from data with missing responses. By using unbiased surrogate estimators, our three-stage procedure avoids imputation while simultaneously performing variable selection and learning the conditional dependence structure among responses. We establish theoretical error bounds, and our simulations demonstrate that \texttt{missoNet} consistently outperforms existing methods in both prediction and sparsity recovery. In a real-world mQTL analysis of the CARTaGENE cohort, \texttt{missoNet} achieved superior predictive accuracy and false-discovery control on a held-out validation set, identifying known and credible novel genetic associations. The method offers a robust, efficient, and theoretically grounded tool for genomic analyses, and is available as an R package.

stat.ME↗