SearcharxivSearch

arXiv subjects

Qingyan Xiang

Publications and source records attributed to Qingyan Xiang.

6 recordsLinked to original sources

Nonlinear trajectories of lung function recovery in patients with pulmonary disease: empirical evaluation of longitudinal modeling approaches

Introduction: Longitudinal lung function recovery after pulmonary disease commonly follows nonlinear trajectories, and failure to adequately model these trajectories can lead to biased or misleading estimates of treatment effects. However, an important methodological gap remains as there is limited assessment of statistical methods for modeling nonlinear lung function trajectories. Methods: We compared several longitudinal modeling approaches for characterizing recovery in percent predicted forced expiratory volume in one second (FEV1p) using data from a phase 2 randomized trial for pulmonary tuberculosis (TB). We estimate the differences in repeated mean FEV1p between each treatment arm and control arm over a 180-days follow-up period. We compared 6 different statistical models: (1) linear mixed-effects model, (2) a piecewise linear mixed-effects model, (3) quadratic and (4) natural cubic spline mixed-effects models, (5) a nonlinear mixed-effects model with exponential recovery function, and (6) a generalized additive mixed model. We discussed the assumptions, clinical interpretations, and resulting treatment-effect estimates across these approaches. Results: The results from the TB trial analyses showed that the conventional linear mixed-effects model provided limited evidence of treatment differences over follow-up, whereas several flexible models identified significant differences during specific periods of recovery. Conclusion: Flexible longitudinal models can complement conventional linear mixed-effects models by revealing treatment differences at certain periods of follow-up that may be obscured by assuming a single linear trend. The choice of nonlinear modeling strategy should be guided by the scientific objective, available data, and the desired balance between clinical interpretability and flexibility.

stat.AP

Analytical and Bootstrap Confidence Intervals of Double Machine Learning: Simulation studies and an application to rural-urban difference in obesity prevalence

Double Machine Learning (DML) is a popular approach for treatment effect estimation in various settings, which allows a wide range of flexible machine learning methods to be used for nuisance parameter estimation while preserving valid inference. In practice, however, applied researchers must choose among many machine learning algorithms for nuisance models, and the impact of this choice on the variance estimation of DML is not well characterized. We conduct a comprehensive simulation study to compare the coverage probability of DML confidence intervals across different machine learning algorithms. In this study, we compare (1) analytical confidence intervals derived by DML theory versus (2) bootstrap confidence interval. We use a set of learners including ordinary least squares, LASSO, Random Forest, LightGBM, and Neural Networks under different data generation settings. We evaluate the performance across difference settings by bias, confidence interval width, and most importantly, coverage probability. Our results show substantial variability in coverage performance across analytical and bootstrap confidence intervals, highlighting that learner choice plays a critical role in reliable DML inference. Surprisingly, we find that in many settings, when sample size increases, the coverage probability of both DML analytical and bootstrap confidence interval decreases. We further investigate coverage probabilities using a real dataset on rural urban differences among U.S. counties. The real data analysis discovers that (1) the model performance still varies by the learner choices and (2) greater rurality has a statistically significant increasing effect on county level obesity prevalence.

stat.ML

Using large language models for sensitivity analysis in causal inference: case studies on Cornfield inequality and E-value

Sensitivity analysis methods such as the Cornfield inequality and the E-value were developed to assess the robustness of observed associations against unmeasured confounding -- a major challenge in observational studies. However, the calculation and interpretation of these methods can be difficult for clinicians and interdisciplinary researchers. Recent advances in large language models (LLMs) offer accessible tools that could assist sensitivity analyses, but their reliability in this context has not been studied. We assess four widely used LLMs, ChatGPT, Claude, DeepSeek, and Gemini, on their ability to conduct sensitivity analyses using Cornfield inequalities and E-values. We first extract study-specific information (exposures, outcomes, measured confounders, and effect estimates) from four published observational studies in different fields. Using such information, we develop structured prompts to assess the performance of the LLMs in three aspects: (1) accuracy of E-value calculation, (2) qualitative interpretation of robustness to unmeasured confounding, and (3) suggestion of possible unmeasured confounders. To our knowledge, there has been little prior work on using LLMs for sensitivity analysis, and this study is an early investigation in this area. The results show that ChatGPT, Claude, and Gemini accurately reproduce the E-values, whereas DeepSeek shows small biases. Qualitative conclusions from all the LLMs align with the magnitude of the E-values and the reported effect sizes, and all models identify biologically and epidemiologically plausible unmeasured confounders. These findings suggest that, when guided by structured prompts, LLMs can effectively assist in evaluating unmeasured confounding, and thereby can support study design and decision-making in observational studies.

stat.OT

Double machine learning to estimate the effects of multiple treatments and their interactions

Causal inference literature has extensively focused on binary treatments, with relatively fewer methods developed for multi-valued treatments. In particular, methods for multiple simultaneously assigned treatments remain understudied despite their practical importance. This paper introduces two settings: (1) estimating the effects of multiple treatments of different types (binary, categorical, and continuous) and the effects of treatment interactions, and (2) estimating the average treatment effect across categories of multi-valued regimens. To obtain robust estimates for both settings, we propose a class of methods based on the Double Machine Learning (DML) framework. Our methods are well-suited for complex settings of multiple treatments/regimens, using machine learning to model confounding relationships while overcoming regularization and overfitting biases through Neyman orthogonality and cross-fitting. To our knowledge, this work is the first to apply machine learning for robust estimation of interaction effects in the presence of multiple treatments. We further establish the asymptotic distribution of our estimators and derive variance estimators for statistical inference. Extensive simulations demonstrate the performance of our methods. Finally, we apply the methods to study the effect of three treatments on HIV-associated kidney disease in an adult HIV cohort of 2455 participants in Nigeria.

stat.ME

Estimating treatment effects from observational data under truncation by death using survival-incorporated quantiles

The issue of "truncation by death" commonly arises in clinical research: subjects may die before their follow-up assessment, resulting in undefined clinical outcomes. To address this issue, we focus on survival-incorporated quantiles -- quantiles of a composite outcome combining death and clinical outcomes -- to summarize the effect of treatment. Using inverse probability of treatment weighting (IPTW), we propose an estimator for survival-incorporated quantiles from observational data, applicable to settings of both point treatment and time-varying treatments. We establish consistency and asymptotic normality of the estimator under both the true and estimated propensity scores. While the variance properties of IPTW estimators for the mean have been studied, to our knowledge, this article is the first to show that the IPTW quantile estimator using the estimated propensity score yields lower asymptotic variance than the IPTW quantile estimator using the true propensity score. Extensive simulations show that survival-incorporated quantiles provide a simple and useful summary measure and confirm that using the estimated propensity score reduces the root mean square error. We apply our method to estimate the effect of statins on the change in cognitive function, incorporating death, using data from the Long Life Family Study (LLFS) -- a multicenter observational study of 4953 older adults with familial longevity. Our results indicate no significant difference in cognitive decline between statin users and non-users with a similar age- and sex-distribution at baseline. This study not only contributes to understand the cognitive effects of statins but also provides insights into analyzing clinical outcomes in the presence of death.

stat.ME

The survival-incorporated median versus the median in the survivors or in the always-survivors: What are we measuring? And why?

Many clinical studies evaluate the benefit of a treatment based on both survival and other continuous/ordinal clinical outcomes, such as Quality of Life scores. In these studies, when subjects die before the follow-up assessment, the clinical outcomes become undefined and are truncated by death. Treating outcomes as "missing" or "censored" due to death can be misleading for treatment effect evaluation. We show that if we use the median in the survivors or in the always-survivors as estimands to summarize clinical outcomes, we may conclude that a trade-off exists between the probability of survival and good clinical outcomes, even in settings where both the probability of survival and the probability of any good clinical outcome are better for one treatment. Therefore, we advocate not always treating death as a mechanism through which clinical outcomes are missing, but rather as part of the outcome measure. To account for the survival status, we describe the survival-incorporated median as an alternative summary measure for outcomes in the presence of death. The survival-incorporated median is the threshold such that 50% of the population is alive with an outcome above that threshold. Through conceptual examples and an application to a prostate cancer treatment study, we show that the survival-incorporated median provides a simple and useful summary measure to inform clinical practice.

stat.AP