SearcharxivSearch

arXiv subjects

Guannan Zhai

Publications and source records attributed to Guannan Zhai.

2 recordsLinked to original sources

Valid test for multi-arm trials with generalized linear models under covariate-adaptive randomization

Modern medical research, such as dose-finding studies, seamless trials, and shared control designs, often involves comparing multiple treatments simultaneously. Despite its wide applications, most research focuses on continuous endpoints, leaving the inference for general outcome types in high demand. In this article, we propose a new inference method for conducting multiple-treatment comparisons involving endpoints within the generalized linear model (GLM) framework under covariate-adaptive randomization (CAR). First, we investigate the asymptotic properties of the standard Wald z-statistics (z-scores) in multi-arm trials, highlighting issues when the working model is misspecified, particularly through omitted covariates. Our theoretical findings reveal that these \textcolor{black}{z-scores} do not consistently converge to a standard multivariate normal distribution, leading to either conservative or inflated Type I error rates, depending on the specific GLM endpoint. Second, based on these theoretical results, we develop adjusted test statistics to correct the distributional problems. To appropriately control the family-wise Type I error rate inherent in multi-arm comparisons, we incorporate our adjusted statistics with Simes-type multiple-testing procedures. This robust inference method can effectively control Type I error while potentially improving power. Extensive simulation studies and a real-world application to a metastatic breast cancer trial confirm the effectiveness and practicality of our approach.

stat.ME

SynthIPD: training-free synthetic individual patient data generation

Individual patient data (IPD) are essential for statistical inference in clinical research. However, privacy concerns, high data-sharing costs, and restrictive access often make IPD unavailable. Conventional synthetic data generation usually relies on black box models such as generative adversial networks. These methods, however, requires a large piece of IPD for model training, may be ungeneralizable and lacks interpretability. This paper introduces an assumption-lean, three-step methodology for generating synthetic IPD with survival endpoints only based on published clinical trial articles. The method mainly leverages Kaplan-Meier (KM) curves with at-risk/censoring information and subgroup-level summary statistics. It digitizes the KM curve using Scalable Vector Graphics (SVG) beyond pixel accuracy and then generates synthetic covariates based on the statistics. We illustrate the method's potential through $2$ detailed case studies and simulation studies. The method offers important implications, enabling high-fidelity IPD generation to support evidence-based medical decisions.

stat.AP