SearcharxivSearch

arXiv subjects

Yongming Qu

Publications and source records attributed to Yongming Qu.

At least 19 recordsLinked to original sources

Empirical Simulation of Survival and Mixed-Type Data for Clinical Trial Design

Simulating realistic time-to-event data is essential for planning and evaluating complex clinical trial designs. Conventional approaches often sample event times from parametric families, such as Weibull or log-normal distributions, which restrict hazard shapes and may poorly represent observed survival data. We propose an empirical copula-based framework for simulating multivariate data containing continuous, binary, count, and right-censored time-to-event variables. The method completes censored historical survival data using a two-zone procedure that combines conditional Kaplan-Meier imputation with a parametric tail. It matches a target survival distribution through a log-scale location-scale transformation and a power distortion of the empirical percentile function, while preserving historical dependence through a Gaussian copula fitted to rank correlations. In an oncology trial of previously treated non-small-cell lung cancer, the method reconstructs overall survival and progression-free survival curves for the experimental arm using control-arm data and a small set of target percentiles. Simulations preserve rank correlations among baseline covariates and the dependence between progression-free and overall survival, with censored Kendall's tau of 0.522 compared with 0.549 in the observed data. The method is implemented in the R package EmpiricalSim.

stat.ME

Optimizing Efficiency and Convergence in MMRM: Practical Considerations for Longitudinal Data Analysis

Mixed models for repeated measures (MMRM) are a popular method for analyzing longitudinal data in clinical trials. However, practical challenges, such as small sample sizes, large numbers of time points, and selection of variance-covariance structure for within-subject errors, often present barriers to model convergence and valid inference. This article evaluates different options for using MMRM in different scenarios through extensive simulation studies and an application on diabetes trial data. We demonstrate that the empirical bias-reduced coefficient covariance adjustment with the heterogeneous autoregressive covariance structure yields near-nominal coverage with high convergence rates for moderate and large sample designs. For small sample sizes, the simple model with baseline covariates and treatment by time point interaction achieves good efficiency and high probability of convergence. Based on these results, we provide practitioners with actionable guidance for applying MMRM to clinical trial data.

stat.ME

An Empirical Method for Analyzing Count Data

Count endpoints are common in clinical trials, particularly for recurrent events such as hypoglycemia. When interest centers on comparing overall event rates between treatment groups, negative binomial (NB) regression is widely used because it accommodates overdispersion and requires only event counts and exposure times. However, NB regression can be numerically unstable when events are sparse, and the efficiency gains from baseline covariate adjustment may be sensitive to model misspecification. We propose an empirical method that targets the same marginal estimand as NB regression -- the ratio of marginal event rates -- while avoiding distributional assumptions on the count outcome. Simulation studies show that the empirical method maintains appropriate Type I error control across diverse scenarios, including extreme overdispersion and zero inflation, achieves power comparable to NB regression, and yields consistent efficiency gains from baseline covariate adjustment. We illustrate the approach using severe hypoglycemia data from the QWINT-5 trial comparing insulin efsitora alfa with insulin degludec in adults with type 1 diabetes. In this sparse-event setting, the empirical method produced stable marginal rate estimates and rate ratios closely aligned with observed rates, while NB regression exhibited greater sensitivity and larger deviations from the observed rates in the sparsest intervals. The proposed empirical method provides a robust and numerically stable alternative to NB regression, particularly when the number of events is low or when numerical stability is a concern.

stat.ME

Estimating treatment effects with competing intercurrent events in randomized controlled trials

The analysis of randomized controlled trials is often complicated by intercurrent events (IEs) -- events that occur after treatment initiation and affect either the interpretation or existence of outcome measurements. Examples include treatment discontinuation or the use of additional medications. In two recent clinical trials for systemic lupus erythematosus with complications of IEs, we classify the IEs into two broad categories: effect-informative (e.g., treatment discontinuation due to adverse events or lack of efficacy) and effect-uninformative (e.g., treatment discontinuation due to external factors such as pandemics or relocation). To define a clinically meaningful estimand, we adopt tailored strategies for each category of IEs. For effect-informative IEs, which are often informative about a patient's outcome, we use the composite variable strategy that assigns an outcome value indicative of treatment failure. For effect-uninformative IEs, we apply the hypothetical strategy, assuming their timing is conditionally independent of the outcome given treatment and baseline covariates, and hypothesizing a scenario in which such events do not occur. A central yet previously overlooked challenge is the presence of competing IEs, where the first IE censors all subsequent ones. Despite its ubiquity in practice, this issue has not been explicitly recognized or addressed in previous data analyses due to the lack of rigorous statistical methodology. In this paper, we propose a principled framework to formulate the estimand, establish its nonparametric identification and semiparametric estimation theory, and introduce weighting, outcome regression, and doubly robust estimators. We apply our methods to analyze the two systemic lupus erythematosus trials, demonstrating the robustness and practical utility of the proposed framework.

stat.ME

Retrieved dropout imputation considering administrative study withdrawal

The International Council for Harmonisation of Technical Requirements for Pharmaceuticals for Human Use (ICH) E9 (R1) Addendum provides a framework for defining estimands in clinical trials. Treatment policy strategy is the mostly used approach to handle intercurrent events in defining estimands. Imputing missing values for potential outcomes under the treatment policy strategy has been discussed in the literature. Missing values as a result of administrative study withdrawals (such as site closures due to business reasons, COVID-19 control measures, and geopolitical conflicts, etc.) are often imputed in the same way as other missing values occurring after intercurrent events related to safety or efficacy. Some research suggests using a hypothetical strategy to handle the treatment discontinuations due to administrative study withdrawal in defining the estimands and imputing the missing values based on completer data assuming missing at random, but this approach ignores the fact that subjects might experience other intercurrent events had they not had the administrative study withdrawal. In this article, we consider the administrative study withdrawal censors the normal real-world like intercurrent events and propose two methods for handling the corresponding missing values under the retrieved dropout imputation framework. Simulation shows the two methods perform well. We also applied the methods to actual clinical trial data evaluating an anti-diabetes treatment.

stat.AP

Direct Estimation for Commonly Used Pattern-Mixture Models in Clinical Trials

Pattern-mixture models have received increasing attention as they are commonly used to assess treatment effects in primary or sensitivity analyses for clinical trials with nonignorable missing data. Pattern-mixture models have traditionally been implemented using multiple imputation, where the variance estimation may be a challenge because the Rubin's approach of combining between- and within-imputation variance may not provide consistent variance estimation while bootstrap methods may be time-consuming. Direct likelihood-based approaches have been proposed in the literature and implemented for some pattern-mixture models, but the assumptions are sometimes restrictive, and the theoretical framework is fragile. In this article, we propose an analytical framework for an efficient direct likelihood estimation method for commonly used pattern-mixture models corresponding to return-to-baseline, jump-to-reference, placebo washout, and retrieved dropout imputations. A parsimonious tipping point analysis is also discussed and implemented. Results from simulation studies demonstrate that the proposed methods provide consistent estimators. We further illustrate the utility of the proposed methods using data from a clinical trial evaluating a treatment for type 2 diabetes.

stat.ME

Implementation of ICH E9 (R1): a few points learned during the COVID-19 pandemic

The current COVID-19 pandemic poses numerous challenges for ongoing clinical trials and provides a stress-testing environment for the existing principles and practice of estimands in clinical trials. The pandemic may increase the rate of intercurrent events (ICEs) and missing values, spurring a great deal of discussion on amending protocols and statistical analysis plans to address these issues. In this article we revisit recent research on estimands and handling of missing values, especially the ICH E9 (R1) on Estimands and Sensitivity Analysis in Clinical Trials. Based on an in-depth discussion of the strategies for handling ICEs using a causal inference framework, we suggest some improvements in applying the estimand and estimation framework in ICH E9 (R1). Specifically, we discuss a mix of strategies allowing us to handle ICEs differentially based on reasons for ICEs. We also suggest ICEs should be handled primarily by hypothetical strategies and provide examples of different hypothetical strategies for different types of ICEs as well as a road map for estimation and sensitivity analyses. We conclude that the proposed framework helps streamline translating clinical objectives into targets of statistical inference and automatically resolves many issues with defining estimands and choosing estimation procedures arising from events such as the pandemic.

stat.ME

Missing data imputation for a multivariate outcome of mixed variable types

Data collected in clinical trials are often composed of multiple types of variables. For example, laboratory measurements and vital signs are longitudinal data of continuous or categorical variables, adverse events may be recurrent events, and death is a time-to-event variable. Missing data due to patients' discontinuation from the study or as a result of handling intercurrent events using a hypothetical strategy almost always occur during any clinical trial. Imputing these data with mixed types of variables simultaneously is a challenge that has not been studied extensively. In this article, we propose using an approximate fully conditional specification to impute the missing data. Simulation shows the proposed method provides satisfactory results under the assumption of missing at random. Finally, real data from a clinical trial evaluating treatments for diabetes are analyzed to illustrate the potential benefit of the proposed method.

stat.ME

Accurate collection of reasons for treatment discontinuation to better define estimands in clinical trials

Background: Reasons for treatment discontinuation are important not only to understand the benefit and risk profile of experimental treatments, but also to help choose appropriate strategies to handle intercurrent events in defining estimands. The current case report form (CRF) commonly in use mixes the underlying reasons for treatment discontinuation and who makes the decision for treatment discontinuation, often resulting in an inaccurate collection of reasons for treatment discontinuation. Methods and results: We systematically reviewed and analyzed treatment discontinuation data from nine phase 2 and phase 3 studies for insulin peglispro. A total of 857 participants with treatment discontinuation were included in the analysis. Our review suggested that, due to the vague multiple-choice options for treatment discontinuation present in the CRF, different reasons were sometimes recorded for the same underlying reason for treatment discontinuation. Based on our review and analysis, we suggest an intermediate solution and a more systematic way to improve the current CRF for treatment discontinuations. Conclusion: This research provides insight and directions on how to optimize the CRF for recording treatment discontinuation. Further work needs to be done to build the learning into Clinical Data Interchange Standards Consortium standards.

stat.AP

Using principal stratification in analysis of clinical trials

The ICH E9(R1) addendum (2019) proposed principal stratification (PS) as one of five strategies for dealing with intercurrent events. Therefore, understanding the strengths, limitations, and assumptions of PS is important for the broad community of clinical trialists. Many approaches have been developed under the general framework of PS in different areas of research, including experimental and observational studies. These diverse applications have utilized a diverse set of tools and assumptions. Thus, need exists to present these approaches in a unifying manner. The goal of this tutorial is threefold. First, we provide a coherent and unifying description of PS. Second, we emphasize that estimation of effects within PS relies on strong assumptions and we thoroughly examine the consequences of these assumptions to understand in which situations certain assumptions are reasonable. Finally, we provide an overview of a variety of key methods for PS analysis and use a real clinical trial example to illustrate them. Examples of code for implementation of some of these approaches are given in supplemental materials.

stat.ME

Assessing the commonly used assumptions in estimating the principal causal effect in clinical trials

In clinical trials, it is often of interest to understand the principal causal effect (PCE), the average treatment effect for a principal stratum (a subset of patients defined by the potential outcomes of one or more post-baseline variables). Commonly used assumptions include monotonicity, principal ignorability, and cross-world assumptions of principal ignorability and principal strata independence. In this article, we evaluate these assumptions through a 2$\times$2 cross-over study in which the potential outcomes under both treatments can be observed, provided there are no carry-over and study period effects. From this example, it seemed the monotonicity assumption and the within-treatment principal ignorability assumptions did not hold well. On the other hand, the assumptions of cross-world principal ignorability and cross-world principal stratum independence conditional on baseline covariates seemed reasonable. With the latter assumptions, we estimated the PCEs, defined by whether the blood glucose standard deviation increased in each treatment period, without relying on the cross-over feature, producing estimates close to the results when exploiting the cross-over feature. To the best of our knowledge, this article is the first attempt to evaluate the plausibility of commonly used assumptions for estimating PCEs using a cross-over trial.

stat.ME

Selection bias in the treatment effect for a principal stratum

Estimation of treatment effect for principal strata has been studied for more than two decades. Existing research exclusively focuses on the estimation, but there is little research on forming and testing hypotheses for principal stratification-based estimands. In this brief report, we discuss a phenomenon in which the true treatment effect for a principal stratum may not equal zero even if the two treatments have the same effect at patient level which implies an equal average treatment effect for the principal stratum. We explain this phenomenon from the perspective of selection bias. This is an important finding and deserves attention when using and interpreting results based on principal stratification. There is a need to further study how to form the null hypothesis for estimands for a principal stratum.

stat.ME

Estimating the treatment effect for adherers using multiple imputation

Randomized controlled trials are considered the gold standard to evaluate the treatment effect (estimand) for efficacy and safety. According to the recent International Council on Harmonisation (ICH)-E9 addendum (R1), intercurrent events (ICEs) need to be considered when defining an estimand, and principal stratum is one of the five strategies to handle ICEs. Qu et al. (2020, Statistics in Biopharmaceutical Research 12:1-18) proposed estimators for the adherer average causal effect (AdACE) for estimating the treatment difference for those who adhere to one or both treatments based on the causal-inference framework, and demonstrated the consistency of those estimators; however, this method requires complex custom programming related to high-dimensional numeric integrations. In this article, we implemented the AdACE estimators using multiple imputation (MI) and constructs CI through bootstrapping. A simulation study showed that the MI-based estimators provided consistent estimators with the nominal coverage probabilities of CIs for the treatment difference for the adherent populations of interest. As an illustrative example, the new method was applied to data from a real clinical trial comparing 2 types of basal insulin for patients with type 1 diabetes.

stat.ME

Return-to-baseline multiple imputation for missing values in clinical trials

Return-to-baseline is an important method to impute missing values or unobserved potential outcomes when certain hypothetical strategies are used to handle intercurrent events in clinical trials. Current return-to-baseline approaches seen in literature and in practice inflate the variability of the "complete" dataset after imputation and lead to biased mean estimators {when the probability of missingness depends on the observed baseline and/or postbaseline intermediate outcomes}. In this article, we first provide a set of criteria a return-to-baseline imputation method should satisfy. Under this framework, we propose a novel return-to-baseline imputation method. Simulations show the completed data after the new imputation approach have the proper distribution, and the estimators based on the new imputation method outperform the traditional method in terms of both bias and variance, when missingness depends on the observed values. The new method can be implemented easily with the existing multiple imputation procedures in commonly used statistical packages.

stat.ME

Analysis of an Incomplete Binary Outcome Dichotomized From an Underlying Continuous Variable in Clinical Trials

In many clinical trials, outcomes of interest include binary-valued endpoints. It is not uncommon that a binary-valued outcome is dichotomized from a continuous outcome at a threshold of clinical interest. To reach the objective, common approaches include (a) fitting the generalized linear mixed model (GLMM) to the dichotomized longitudinal binary outcome and (b) imputation method (MI): imputing the missing values in the continuous outcome, dichotomizing it into a binary outcome, and then fitting the generalized linear model for the "complete" data. We conducted comprehensive simulation studies to compare the performance of GLMM with MI for estimating risk difference and logarithm of odds ratio between two treatment arms at the end of study. In those simulation studies, we considered a range of multivariate distribution options for the continuous outcome (including a multivariate normal distribution, a multivariate t-distribution, a multivariate log-normal distribution, and the empirical distribution from a real clinical trial data) to evaluate the robustness of the estimators to various data-generating models. Simulation results demonstrate that both methods work well under those considered distribution options, but MI is more efficient with smaller mean squared errors compared to GLMM. We further applied both the GLMM and MI to 29 phase 3 diabetes clinical trials, and found that the MI method generally led to smaller variance estimates compared to GLMM.

stat.AP

Understanding and adjusting the selection bias from a proof-of-concept study to a more confirmatory study

It has long been noticed that the efficacy observed in small early phase studies is generally better than that observed in later larger studies. Historically, the inflation of the efficacy results from early proof-of-concept studies is either ignored, or adjusted empirically using a frequentist or Bayesian approach. In this article, we systematically explained the underlying reason for the inflation of efficacy results in small early phase studies from the perspectives of measurement error models and selection bias. A systematic method was built to adjust the early phase study results from both frequentist and Bayesian perspectives. A hierarchical model was proposed to estimate the distribution of the efficacy for a portfolio of compounds, which can serve as the prior distribution for the Bayesian approach. We showed through theory that the systematic adjustment provides an unbiased estimator for the true mean efficacy for a portfolio of compounds. The adjustment was applied to paired data for the efficacy in early small and later larger studies for a set of compounds in diabetes and immunology. After the adjustment, the bias in the early phase small studies seems to be diminished.

stat.ME

Defining Estimands Using a Mix of Strategies to Handle Intercurrent Events in Clinical Trials

Randomized controlled trials (RCT) are the gold standard for evaluation of the efficacy and safety of investigational interventions. If every patient in an RCT were to adhere to the randomized treatment, one could simply analyze the complete data to infer the treatment effect. However, intercurrent events (ICEs) including the use of concomitant medication for unsatisfactory efficacy, treatment discontinuation due to adverse events, or lack of efficacy, may lead to interventions that deviate from the original treatment assignment. Therefore, defining the appropriate estimand (the appropriate parameter to be estimated) based on the primary objective of the study is critical prior to determining the statistical analysis method and analyzing the data. The International Council for Harmonisation (ICH) E9 (R1), published on November 20, 2019, provided 5 strategies to define the estimand: treatment policy, hypothetical, composite variable, while on treatment and principal stratum. In this article, we propose an estimand using a mix of strategies in handling ICEs. This estimand is an average of the null treatment difference for those with ICEs potentially related to safety and the treatment difference for the other patients if they would complete the assigned treatments. Two examples from clinical trials evaluating anti-diabetes treatments are provided to illustrate the estimation of this proposed estimand and to compare it with the estimates for estimands using hypothetical and treatment policy strategies in handling ICEs.

stat.AP

Implementation of Tripartite Estimands Using Adherence Causal Estimators Under the Causal Inference Framework

Intercurrent events (ICEs) and missing values are inevitable in clinical trials of any size and duration, making it difficult to assess the treatment effect for all patients in randomized clinical trials. Defining the appropriate estimand that is relevant to the clinical research question is the first step in analyzing data. The tripartite estimands, which evaluate the treatment differences in the proportion of patients with ICEs due to adverse events, the proportion of patients with ICEs due to lack of efficacy, and the primary efficacy outcome for those who can adhere to study treatment under the causal inference framework, are of interest to many stakeholders in understanding the totality of treatment effects. In this manuscript, we discuss the details of how to estimate tripartite estimands based on a causal inference framework and how to interpret tripartite estimates through a phase 3 clinical study evaluating a basal insulin treatment for patients with type 1 diabetes.

stat.AP