SearcharxivSearch

arXiv subjects

Feifang Hu

Publications and source records attributed to Feifang Hu.

At least 19 recordsLinked to original sources

Discretization in covariate-adaptive randomization: gains and losses

Covariate-adaptive randomization(CAR) is widely implemented in clinical trials to balance prognostic covariates across treatment arms. Continuous covariates are often discretized into strata in practice, yet their consequences are not clearly understood. This paper provides a comprehensive study of the impact of discretization on both the CAR design process and the inferential results thereafter. We establish the asymptotic properties of both imbalance measures and treatment effect estimators under discretized and non-discretized settings. Practical recommendations are given on when and how discretization should be employed. We show that discretization in design is generally recommended, as it enhances robustness against model misspecification. However, if the true model is known, the most efficient strategy is to balance covariates according to that model in the design. The theoretical results are corroborated by extensive simulation studies and an empirical application to a diabetes trial dataset. Together, the results clarify the gains and losses of discretization in CAR and pave the way for learning impact of discretization to other designs and beyond.

stat.ME

Valid test for multi-arm trials with generalized linear models under covariate-adaptive randomization

Modern medical research, such as dose-finding studies, seamless trials, and shared control designs, often involves comparing multiple treatments simultaneously. Despite its wide applications, most research focuses on continuous endpoints, leaving the inference for general outcome types in high demand. In this article, we propose a new inference method for conducting multiple-treatment comparisons involving endpoints within the generalized linear model (GLM) framework under covariate-adaptive randomization (CAR). First, we investigate the asymptotic properties of the standard Wald z-statistics (z-scores) in multi-arm trials, highlighting issues when the working model is misspecified, particularly through omitted covariates. Our theoretical findings reveal that these \textcolor{black}{z-scores} do not consistently converge to a standard multivariate normal distribution, leading to either conservative or inflated Type I error rates, depending on the specific GLM endpoint. Second, based on these theoretical results, we develop adjusted test statistics to correct the distributional problems. To appropriately control the family-wise Type I error rate inherent in multi-arm comparisons, we incorporate our adjusted statistics with Simes-type multiple-testing procedures. This robust inference method can effectively control Type I error while potentially improving power. Extensive simulation studies and a real-world application to a metastatic breast cancer trial confirm the effectiveness and practicality of our approach.

stat.ME

Systematic Literature Reviews With Two Multi-Agentic Systems And Human-In-The-Loop

Systematic literature review of clinical trials drives regulatory decision-making, but conventional screening and extraction are time-consuming, labor-intensive, and vulnerable to study selection bias. We propose two fit-to-purpose multi-agentic systems (MAS) for systematic literature review, with human-in-the-loop. The screening MAS uses multiple LLM agents with heterogeneous personas and multiround cross-review, and uniformly improves accuracy over a single-LLM baseline. The extraction MAS combines standardization, an iterative correction loop, and retrieval-based context control to ensure accuracy and scalability. Both MAS are specifically designed to support Human-In-The-Loop which is essential for clinical decisions. The novelty of the proposed approach lies in the system architecture rather than in any single foundation tools: the system can naturally benefit from future improvements in the underlying tools, for instance, stronger LLM agents, retrieval engines, image recognition methods, etc. As a real-world application, a published network meta-analysis is reproduced by the MAS. The result recovers all trials from the original study and identifies additional eligible trials missed by manual review, leading to updated clinical conclusions.

stat.AP

Estimating Treatment and Spillover Effects with the Ego-Cluster Experimental Design

Network interference occurs when a unit's outcome depends not only on its own treatment but also on the treatments received by connected units in the network. Experimental designs and analysis methods that ignore such interference can yield biased estimators of causal effects. In this paper, we develop a new experimental design for the estimation and inference of global treatment effect and spillover effect under a model-based framework and ego-cluster randomization. Under this design, the network is partitioned into a collection of ego-clusters, each consisting of a focal unit (the ego) and its network neighbors (the alters), with randomization conducted at the cluster level. We propose model-based estimators for the global treatment effect and spillover effect and establish their consistency and asymptotic normality, with asymptotic variances determined by the ego-cluster structure. Building on these theoretical results, we introduce an ego-clustering algorithm that sequentially selects egos and assigns alters to minimize asymptotic variances. Simulation studies and two empirical applications demonstrate that the proposed procedure yields accurate inference and efficiency improvements over existing network experimental designs.

stat.ME

SynthIPD: training-free synthetic individual patient data generation

Individual patient data (IPD) are essential for statistical inference in clinical research. However, privacy concerns, high data-sharing costs, and restrictive access often make IPD unavailable. Conventional synthetic data generation usually relies on black box models such as generative adversial networks. These methods, however, requires a large piece of IPD for model training, may be ungeneralizable and lacks interpretability. This paper introduces an assumption-lean, three-step methodology for generating synthetic IPD with survival endpoints only based on published clinical trial articles. The method mainly leverages Kaplan-Meier (KM) curves with at-risk/censoring information and subgroup-level summary statistics. It digitizes the KM curve using Scalable Vector Graphics (SVG) beyond pixel accuracy and then generates synthetic covariates based on the statistics. We illustrate the method's potential through $2$ detailed case studies and simulation studies. The method offers important implications, enabling high-fidelity IPD generation to support evidence-based medical decisions.

stat.AP

Testing for Treatment Effect in Covariate-Adaptive Randomized Clinical Trials with Generalized Linear Models and Omitted Covariates

Concerns have been expressed over the validity of statistical inference under covariate-adaptive randomization despite the extensive use in clinical trials. In the literature, the inferential properties under covariate-adaptive randomization have been mainly studied for continuous responses; in particular, it is well known that the usual two sample t-test for treatment effect is typically conservative, in the sense that the actual test size is smaller than the nominal level. This phenomenon of invalid tests has also been found for generalized linear models without adjusting for the covariates and are sometimes more worrisome due to inflated Type I error. The purpose of this study is to examine the unadjusted test for treatment effect under generalized linear models and covariate-adaptive randomization. For a large class of covariate-adaptive randomization methods, we obtain the asymptotic distribution of the test statistic under the null hypothesis and derive the conditions under which the test is conservative, valid, or anti-conservative. Several commonly used generalized linear models, such as logistic regression and Poisson regression, are discussed in detail. An adjustment method is also proposed to achieve a valid size based on the asymptotic results. Numerical studies confirm the theoretical findings and demonstrate the effectiveness of the proposed adjustment method.

stat.ME

Statistical Inference for Covariate-Adaptive Randomization Procedures

Covariate-adaptive randomization (CAR) procedures are frequently used in comparative studies to increase the covariate balance across treatment groups. However, because randomization inevitably uses the covariate information when forming balanced treatment groups, the validity of classical statistical methods after such randomization is often unclear. In this article, we derive the theoretical properties of statistical methods based on general CAR under the linear model framework. More importantly, we explicitly unveil the relationship between covariate-adaptive and inference properties by deriving the asymptotic representations of the corresponding estimators. We apply the proposed general theory to various randomization procedures such as complete randomization, rerandomization, pairwise sequential randomization, and Atkinson's $D_A$-biased coin design and compare their performance analytically. Based on the theoretical results, we then propose a new approach to obtain valid and more powerful tests. These results open a door to understand and analyze experiments based on CAR. Simulation studies provide further evidence of the advantages of the proposed framework and the theoretical results. Supplementary materials for this article are available online.

math.ST

Adaptive Randomization in Network Data

Network data have appeared frequently in recent research. For example, in comparing the effects of different types of treatment, network models have been proposed to improve the quality of estimation and hypothesis testing. In this paper, we focus on efficiently estimating the average treatment effect using an adaptive randomization procedure in networks. We work on models of causal frameworks, for which the treatment outcome of a subject is affected by its own covariate as well as those of its neighbors. Moreover, we consider the case in which, when we assign treatments to the current subject, only the subnetwork of existing subjects is revealed. New randomized procedures are proposed to minimize the mean squared error of the estimated differences between treatment effects. In network data, it is usually difficult to obtain theoretical properties because the numbers of nodes and connections increase simultaneously. Under mild assumptions, our proposed procedure is closely related to a time-varying inhomogeneous Markov chain. We then use Lyapunov functions to derive the theoretical properties of the proposed procedures. The advantages of the proposed procedures are also demonstrated by extensive simulations and experiments on real network data.

stat.ME

Cluster-Adaptive Network A/B Testing: From Randomization to Estimation

A/B testing is an important decision-making tool in product development for evaluating user engagement or satisfaction from a new service, feature or product. The goal of A/B testing is to estimate the average treatment effects (ATE) of a new change, which becomes complicated when users are interacting. When the important assumption of A/B testing, the Stable Unit Treatment Value Assumption (SUTVA), which states that each individual's response is affected by their own treatment only, is not valid, the classical estimate of the ATE usually leads to a wrong conclusion. In this paper, we propose a cluster-adaptive network A/B testing procedure, which involves a sequential cluster-adaptive randomization and a cluster-adjusted estimator. The cluster-adaptive randomization is employed to minimize the cluster-level Mahalanobis distance within the two treatment groups, so that the variance of the estimate of the ATE can be reduced. In addition, the cluster-adjusted estimator is used to eliminate the bias caused by network interference, resulting in a consistent estimation for the ATE. Numerical studies suggest our cluster-adaptive network A/B testing achieves consistent estimation with higher efficiency. An empirical study is conducted based on a real world network to illustrate how our method can benefit decision-making in application.

stat.ME

On the Theory of Covariate-Adaptive Designs

Pocock and Simon's marginal procedure (Pocock and Simon, 1975) is often implemented forbalancing treatment allocation over influential covariates in clinical trials. However, the theoretical properties of Pocock and Simion's procedure have remained largely elusive for decades. In this paper, we propose a general framework for covariate-adaptive designs and establish the corresponding theory under widely satisfied conditions. As a special case, we obtain the theoretical properties of Pocock and Simon's marginal procedure: the marginal imbalances and overall imbalance are bounded in probability, but the within-stratum imbalances increase with the rate of $\sqrt{n}$ as the sample size increases. The theoretical results provide new insights about balance properties of covariate-adaptive randomization procedures and open a door to study the theoretical properties of statistical inference for clinical trials based on covariate-adaptive randomization procedures.

math.ST

Pairwise Sequential Randomization and Its Properties

In comparative studies, such as in causal inference and clinical trials, balancing important covariates is often one of the most important concerns for both efficient and credible comparison. However, chance imbalance still exists in many randomized experiments. This phenomenon of covariate imbalance becomes much more serious as the number of covariates $p$ increases. To address this issue, we introduce a new randomization procedure, called pairwise sequential randomization (PSR). The proposed method allocates the units sequentially and adaptively, using information on the current level of imbalance and the incoming unit's covariate. With a large number of covariates or a large number of units, the proposed method shows substantial advantages over the traditional methods in terms of the covariate balance, estimation accuracy, and computational time, making it an ideal technique in the era of big data. The proposed method attains the optimal covariate balance, in the sense that the estimated treatment effect under the proposed method attains its minimum variance asymptotically. Also the proposed method is widely applicable in both causal inference and clinical trials. Numerical studies and real data analysis provide further evidence of the advantages of the proposed method.

stat.ME

Two-Way Partial AUC and Its Properties

When people evaluate the performance of a diagnostic test, it is important to control both True Positive Rate (TPR) and False Positive Rate (FPR). In the literature, most researchers propose the partial area under the ROC curve (pAUC) with restrictions on FPR to assess a binary classification system, which is named as FPR pAUC. It could be artificially designed to measure the area controlled by TPR and FPR, but is often misleading conceptually and practically. A new and intuitive method, named two-way pAUC, is provided in this paper, which focuses directly on the partial area under the ROC curve with both horizontal and vertical restrictions. We propose a nonparametric estimator of two-way pAUC, obtain its asymptotic normality properties and conduct the measure comparison by bootstrap method. Further, in order to evaluate possible covariate effects on two-way pAUC, regression analysis framework is constructed and corresponding theoretical properties are established. Simulation and real application are conducted to support our methods.

stat.ME

Asymptotic properties of covariate-adaptive randomization

Balancing treatment allocation for influential covariates is critical in clinical trials. This has become increasingly important as more and more biomarkers are found to be associated with different diseases in translational research (genomics, proteomics and metabolomics). Stratified permuted block randomization and minimization methods [Pocock and Simon Biometrics 31 (1975) 103-115, etc.] are the two most popular approaches in practice. However, stratified permuted block randomization fails to achieve good overall balance when the number of strata is large, whereas traditional minimization methods also suffer from the potential drawback of large within-stratum imbalances. Moreover, the theoretical bases of minimization methods remain largely elusive. In this paper, we propose a new covariate-adaptive design that is able to control various types of imbalances. We show that the joint process of within-stratum imbalances is a positive recurrent Markov chain under certain conditions. Therefore, this new procedure yields more balanced allocation. The advantages of the proposed procedure are also demonstrated by extensive simulation studies. Our work provides a theoretical tool for future research in this area.

math.ST

Immigrated urn models - asymptotic properties and applications

Urn models have been widely studied and applied in both scientific and social science disciplines. In clinical studies, the adoption of urn models in treatment allocation schemes has been proved to be beneficial to both researchers, by providing more efficient clinical trials, and patients, by increasing the likelihood of receiving the better treatment. In this paper, we propose a new and general class of immigrated urn (IMU) models that incorporates the immigration mechanism into the urn process. Theoretical properties are developed and the advantages of the IMU models are discussed. In general, the IMU models have smaller variabilities than the classical urn models, yielding more powerful statistical inferences in applications. Illustrative examples are presented to demonstrate the wide applicability of the IMU models. The proposed IMU framework, including many popular classical urn models, not only offers a unify perspective for us to comprehend the urn process, but also enables us to generate several novel urn models with desirable properties.

math.ST

Sequential monitoring of response-adaptive randomized clinical trials

Clinical trials are complex and usually involve multiple objectives such as controlling type I error rate, increasing power to detect treatment difference, assigning more patients to better treatment, and more. In literature, both response-adaptive randomization (RAR) procedures (by changing randomization procedure sequentially) and sequential monitoring (by changing analysis procedure sequentially) have been proposed to achieve these objectives to some degree. In this paper, we propose to sequentially monitor response-adaptive randomized clinical trial and study it's properties. We prove that the sequential test statistics of the new procedure converge to a Brownian motion in distribution. Further, we show that the sequential test statistics asymptotically satisfy the canonical joint distribution defined in Jennison and Turnbull (\citeyearJT00). Therefore, type I error and other objectives can be achieved theoretically by selecting appropriate boundaries. These results open a door to sequentially monitor response-adaptive randomized clinical trials in practice. We can also observe from the simulation studies that, the proposed procedure brings together the advantages of both techniques, in dealing with power, total sample size and total failure numbers, while keeps the type I error. In addition, we illustrate the characteristics of the proposed procedure by redesigning a well-known clinical trial of maternal-infant HIV transmission.

math.ST

Efficient randomized-adaptive designs

Response-adaptive randomization has recently attracted a lot of attention in the literature. In this paper, we propose a new and simple family of response-adaptive randomization procedures that attain the Cramer--Rao lower bounds on the allocation variances for any allocation proportions, including optimal allocation proportions. The allocation probability functions of proposed procedures are discontinuous. The existing large sample theory for adaptive designs relies on Taylor expansions of the allocation probability functions, which do not apply to nondifferentiable cases. In the present paper, we study stopping times of stochastic processes to establish the asymptotic efficiency results. Furthermore, we demonstrate our proposal through examples, simulations and a discussion on the relationship with earlier works, including Efron's biased coin design.

math.ST

The Gaussian approximation for multi-color generalized Friedman's urn model

The Friedman's urn model is a popular urn model which is widely used in many disciplines. In particular, it is extensively used in treatment allocation schemes in clinical trials. In this paper, we prove that both the urn composition process and the allocation proportion process can be approximated by a multi-dimensional Gaussian process almost surely for a multi-color generalized Friedman's urn model with non-homogeneous generating matrices. The Gaussian process is a solution of a stochastic differential equation. This Gaussian approximation together with the properties of the Gaussian process is important for the understanding of the behavior of the urn process and is also useful for statistical inferences. As an application, we obtain the asymptotic properties including the asymptotic normality and the law of the iterated logarithm for a multi-color generalized Friedman's urn model as well as the randomized-play-the-winner rule as a special case.

math.PR