Searcharxiv⌕ Search

arXiv subjects

Xiao-Hua Zhou

Publications and source records attributed to Xiao-Hua Zhou.

At least 19 recordsLinked to original sources

Design-based Estimation and Inference on Quantile Exposure Effect under General Interference

Many applications in public health, environmental science, and economics feature spillovers across connected units, violating the Stable Unit Treatment Value Assumption (SUTVA) underlying classical quantile treatment effect methods. We develop a general framework for defining, estimating, and conducting inference for quantile exposure effects (QEEs) under network interference, encompassing quantile direct and spillover effects as leading cases. Studying QEEs under interference faces three substantive challenges. First, because a single exposure level represents many neighborhood treatment configurations, the causal estimand must aggregate over these configurations in an interpretable manner. Second, quantile estimands are intrinsically non-smooth, placing them outside much of the existing network theory involving only Lipschitz or differentiable functionals. Third, in design-based, finite-population network settings, conventional smoothness assumptions on outcome densities cannot be imposed directly. Using conditional neighborhood dependence, we establish asymptotic normality under explicit network degree conditions and derive a locally uniform convergence rate for kernel density estimation. We also characterize the variance estimation bias arising from heterogeneous unit-specific score means and construct asymptotically conservative confidence intervals. Extensive simulations and an application to an educational intervention in school friendship networks demonstrate that the proposed framework reveals heterogeneous exposure effects, including tail-specific impacts, that are missed by analyses based solely on average treatment effects.

stat.ME↗

Estimating Pathway Treatment Effects in the Presence of Intermediate Events with Multi-State Data

During clinical trials evaluating a drug's effect on a survival endpoint, intermediate events often occur in addition to the primary event. The treatment can exert its effect on the primary endpoint along multiple pathways through intermediate events. Assumptions for identifying mediation effects, such as sequential ignorability in natural effects or the dismissible components condition in separable effects, fail because intermediate events act as treatment-induced confounding. To understand the effect along each pathway, we consider hypothetical interventions in transitions between event statuses to mimic the treatment mechanism. The hypothetical interventions adjust for effects through intermediate events and marginalize over unobserved treatment-induced confounding, if any. Based on the derived efficient influence functions for the counterfactual cumulative incidences under hypothetical interventions, we construct multiply robust and semiparametrically efficient estimators for pathway treatment effects. Our proposed framework enables the examination of treatment effects through each transition, on each event, and along each path. By analyzing data from the LEADER Trial, we find that liraglutide significantly reduces the risk of cardiovascular and microvascular events. The reduction in all-cause mortality is primarily mediated by its effects on expanded major adverse cardiovascular events.

stat.ME↗

A new non-parametric test for multivariate paired data from pair matching or paired designs

In observational studies, achieving covariate balance in pair matching between treatment and control groups or exposed and unexposed groups is essential. This balance enables testing treatment effects or examining {associations between exposures and} multivariate response variables in pair-matched data. Paired design studies involve taking multiple measurements for the same subjects under different conditions. All these call for an effective test for multivariate paired data. However, current methods for assessing covariate balance in matched observational studies often ignore the paired structure, leading to reduced performance in some cases. The multivariate paired Hotelling's $T^2$ test can be used for paired data, but its power decreases rapidly as dimensions increase. To address these issues, we propose a new non-parametric test for paired data, significantly improving power across various scenarios. We also derive the test's asymptotic distribution, making it user-friendly for practical applications. Our proposed test's effectiveness is demonstrated through an analysis of real data on Alzheimer's disease research.

stat.ME↗

A Unified Inference Method for FROC-type Curves and Related Summary Indices

Free-response observer performance studies are of great importance for accuracy evaluation and comparison in tasks related to the detection and localization of multiple targets or signals. The free-response receiver operating characteristic (FROC) curve and many similar curves based on the free-response observer performance assessment data are important tools to display the accuracy of detection under different thresholds. The true positive rate at a fixed false positive rate and summary indices such as the area under the FROC curve are also commonly used as the figures of merit in the statistical evaluation of these studies. Motivated by a free-response observer performance assessment research of a Software as a Medical Device (SaMD), we propose a unified method based on the initial-detection-and-candidate model to simultaneously estimate a smooth curve and derive confidence intervals for summary indices and the true positive rate at a fixed false positive rate. A maximum likelihood estimator is proposed and its asymptotic normality property is derived. Confidence intervals are constructed based on the asymptotic normality of our maximum likelihood estimator. Simulation studies are conducted to evaluate the finite sample performance of the proposed method. We apply the proposed method to evaluate the diagnostic performance of the SaMD for detecting pulmonary lesions.

stat.ME↗

Biomarkers selection and combination based on the weighted Youden index

In clinical practice, multiple biomarkers are used for disease diagnosis, but their individual accuracies are often suboptimal, with only a few proving directly relevant. Effectively selecting and combining biomarkers can significantly improve diagnostic accuracy. Existing methods often optimize metrics like the Area Under the ROC Curve (AUC) or the Youden index. However, optimizing AUC does not yield estimates for optimal cutoff values, and the Youden index assumes equal weighting of sensitivity and specificity, which may not reflect clinical priorities where these metrics are weighted differently. This highlights the need for methods that can flexibly accommodate such requirements. In this paper, we present a novel framework for selecting and combining biomarkers to maximize a weighted version of the Youden index. We introduce a smoothed estimator based on the weighted Youden index and propose a penalized version using the SCAD penalty to enhance variable selection. To handle the non-convexity of the objective function and the non-smoothness of the penalty, we develop an efficient algorithm, also applicable to other non-convex optimization problems. Simulation studies demonstrate the performance and efficiency of our method, and we apply it to construct a diagnostic scale for dermatitis.

stat.ME↗

Copas-Jackson-type bounds for publication bias over a general class of selection models

Publication bias (PB) is one of the most vital threats to the accuracy of meta-analysis. Adjustment or sensitivity analysis based on selection models, which describe the probability of a study being published, provide a more objective evaluation of PB than widely-used simple graphical methods such as the trim-and-fill method. Most existing methods rely on parametric selection models. The Copas-Jackson bound (C-J bound) provides a worst-case bound of an analytical form over a nonparametric class of selection models, which would provide more robust conclusions than parametric sensitivity analysis. The nonparametric class of the selection models in the C-J bound is restrictive and only covers parametric selection models monotonic to the standard errors of outcomes. The novelty of this paper is to develop a method that constructs worst-case bounds over a general class of selection models weakening the assumption in the C-J bound. We propose an efficient numerical method to obtain an approximate worst-case bound via tractable nonlinear programming with linear constraints. We substantiate the effectiveness of the proposed bound with extensive simulation studies and show its applicability with two real-world meta-analyses.

stat.ME↗

Copas-Heckman-type sensitivity analysis for publication bias in rare-event meta-analysis under generalized linear mixed models

In systematic reviews and meta-analyses, publication bias (PB) is one of the serious concerns and mainly induced by selective publication of academic literatures. Although many methods have been proposed to deal with PB, almost all the methods are based on the normal-normal (NN) random-effects model assuming that data are normally distributed in both the within-study and the between-study levels. For rare-event meta-analysis where data contain rare occurrences of events, the standard NN random-effects model may perform poorly. Instead, some generalized linear mixed models (GLMMs) which employ the exact distribution for the number of events in within-study level provide alternatives and have been widely used in practice. However, limited methods can be applied to deal with PB in the GLMMs. To address this limitation, we propose a framework of sensitivity analysis for evaluating the impact of PB in various GLMMs. The proposed framework is developed based on the famous Copas-Heckman-type sensitivity analysis methods and can be easily implemented with the standard software with small computational cost. In this paper, we conduct simulation studies to assess the performance of proposed methods in adjusting PB and compare the results with related existing methods. Several real-world examples are also analyzed to show the broad applicability of our proposal in evaluating the potential impact of PB in meta-analysis of odds ratios and proportions with rare-event outcomes.

stat.ME↗

CSTEapp: An interactive R-Shiny application of the covariate-specific treatment effect curve for visualizing individualized treatment rule

In precision medicine, deriving the individualized treatment rule (ITR) is crucial for recommending the optimal treatment based on patients' baseline covariates. The covariate-specific treatment effect (CSTE) curve presents a graphical method to visualize an ITR within a causal inference framework. Recent advancements have enhanced the causal interpretation of the CSTE curves and provided methods for deriving simultaneous confidence bands for various study types. To facilitate the implementation of these methods and make ITR estimation more accessible, we developed CSTEapp, a web-based application built on the R Shiny framework. CSTEapp allows users to upload data and create CSTE curves through simple point and click operations, making it the first application for estimating the ITRs. CSTEapp simplifies the analytical process by providing interactive graphical user interfaces with dynamic results, enabling users to easily report optimal treatments for individual patients based on their covariates information. Currently, CSTEapp is applicable to studies with binary and time-to-event outcomes, and we continually expand its capabilities to accommodate other outcome types as new methods emerge. We demonstrate the utility of CSTEapp using real-world examples and simulation datasets. By making advanced statistical methods more accessible, CSTEapp empowers researchers and practitioners across various fields to advance precision medicine and improve patient outcomes.

stat.CO↗

Inference for Cumulative Incidences and Treatment Effects in Randomized Controlled Trials with Time-to-Event Outcomes under ICH E9 (R1)

In randomized controlled trials (RCTs) that focus on time-to-event outcomes, intercurrent events can arise in two ways: as semi-competing events, which modify the hazard of the primary outcome events, or as competing events, which make the definition of the primary outcome events unclear. Although five strategies have been proposed in the ICH E9 (R1) addendum to address intercurrent events in RCTs, these strategies are not easily applicable to time-to-event outcomes when aiming for causal interpretations. In this study, we show how to define, estimate, and make inferences concerning objectives that have causal interpretations within these contexts. Specifically, we derive the mathematical formulations of the causal estimands corresponding to the five strategies and clarify the data structure needed to identify these causal estimands. Furthermore, we introduce nonparametric methods for estimating and making inferences about these causal estimands, including the asymptotic variance of estimators and hypothesis tests. Finally, we illustrate our methods using data from the LEADER Trial, which aims to investigate the effect of liraglutide on cardiovascular outcomes.

stat.ME↗

Biomarker combination based on the Youden index with and without gold standard

In clinical practice, multiple biomarkers are often measured on the same subject for disease diagnosis, and combining them can improve diagnostic accuracy. Existing studies typically combine multiple biomarkers by maximizing the Area Under the ROC Curve (AUC), assuming a gold standard exists or that biomarkers follow a multivariate normal distribution. However, practical diagnostic settings require both optimal combination coefficients and an effective cutoff value, and the reference test may be imperfect. In this paper, we propose a two-stage method for identifying the optimal linear combination and cutoff value based on the Youden index. First, it maximizes an approximation of the empirical AUC to estimate the optimal linear coefficients for combining multiple biomarkers. Then, it maximizes the empirical Youden index to determine the optimal cutoff point for disease classification. Under the semiparametric single index model and regularity conditions, the estimators for the linear coefficients, cutoff point, and Youden index are consistent. This method is also applicable when the reference standard is imperfect. We demonstrate the performance of our method through simulations and apply it to construct a diagnostic scale for Chinese medicine.

stat.ME↗

Identifying average causal effect in regression discontinuity design with auxiliary data

Regression discontinuity designs are widely used when treatment assignment is determined by whether a running variable exceeds a predefined threshold. However, most research focuses on estimating local causal effects at the threshold, leaving the challenge of identifying treatment effects away from the cutoff largely unaddressed. The primary difficulty in this context is that the treatment assignment is deterministically defined by the running variable, violating the commonly assumed positivity assumption. In this paper, we introduce a novel framework for identifying the average causal effect in regression discontinuity designs. Our approach assumes the existence of an auxiliary variable for which the running variable can be seen as a surrogate, and an additional dataset that consists of the running variable and the auxiliary variable alongside the traditional regression discontinuity design setup. Under this framework, we propose three estimation methods for the ATE, which resembles the outcome regression, inverse propensity weighted and doubly robust estimators in classical causal inference literature. Asymptotically valid inference procedures are also provided. To demonstrate the practical application of our method, simulations are conducted to show the good performance of our methods; besides, we use the proposed methods to assess the causal effects of vitamin A supplementation on the severity of autism spectrum disorders in children, where a positive effect is found but with no statistical significance.

stat.ME↗

Separable pathway effects of semi-competing risks using multi-state models

Semi-competing risks refer to the phenomenon where a primary event (such as mortality) can ``censor'' an intermediate event (such as relapse of a disease), but not vice versa. Under the multi-state model, the primary event consists of two specific types: the direct outcome event and an indirect outcome event developed from intermediate events. Within this framework, we show that the total treatment effect on the cumulative incidence of the primary event can be decomposed into three separable pathway effects, capturing treatment effects on population-level transition rates between states. We next propose two estimators for the counterfactual cumulative incidences of the primary event under hypothetical treatment components. One estimator is given by the generalized Nelson--Aalen estimator with inverse probability weighting under covariates isolation, and the other is given based on the efficient influence function. The asymptotic normality of these estimators is established. The first estimator only involves a propensity score model and avoid modeling the cause-specific hazards. The second estimator has robustness against the misspecification of submodels. As an illustration of its potential usefulness, the proposed method is applied to compare effects of different allogeneic stem cell transplantation types on overall survival after transplantation.

stat.ME↗

Debiased Recommendation with Noisy Feedback

Ratings of a user to most items in recommender systems are usually missing not at random (MNAR), largely because users are free to choose which items to rate. To achieve unbiased learning of the prediction model under MNAR data, three typical solutions have been proposed, including error-imputation-based (EIB), inverse-propensity-scoring (IPS), and doubly robust (DR) methods. However, these methods ignore an alternative form of bias caused by the inconsistency between the observed ratings and the users' true preferences, also known as noisy feedback or outcome measurement errors (OME), e.g., due to public opinion or low-quality data collection process. In this work, we study intersectional threats to the unbiased learning of the prediction model from data MNAR and OME in the collected data. First, we design OME-EIB, OME-IPS, and OME-DR estimators, which largely extend the existing estimators to combat OME in real-world recommendation scenarios. Next, we theoretically prove the unbiasedness and generalization bound of the proposed estimators. We further propose an alternate denoising training approach to achieve unbiased learning of the prediction model under MNAR data with OME. Extensive experiments are conducted on three real-world datasets and one semi-synthetic dataset to show the effectiveness of our proposed approaches. The code is available at https://github.com/haoxuanli-pku/KDD24-OME-DR.

cs.IR↗

Phased Instruction Fine-Tuning for Large Language Models

Instruction Fine-Tuning enhances pre-trained language models from basic next-word prediction to complex instruction-following. However, existing One-off Instruction Fine-Tuning (One-off IFT) method, applied on a diverse instruction, may not effectively boost models' adherence to instructions due to the simultaneous handling of varying instruction complexities. To improve this, Phased Instruction Fine-Tuning (Phased IFT) is proposed, based on the idea that learning to follow instructions is a gradual process. It assesses instruction difficulty using GPT-4, divides the instruction data into subsets of increasing difficulty, and uptrains the model sequentially on these subsets. Experiments with Llama-2 7B/13B/70B, Llama3 8/70B and Mistral-7B models using Alpaca data show that Phased IFT significantly outperforms One-off IFT, supporting the progressive alignment hypothesis and providing a simple and efficient way to enhance large language models. Codes and datasets from our experiments are freely available at https://github.com/xubuvd/PhasedSFT.

cs.CL↗

A likelihood-based sensitivity analysis for addressing publication bias in meta-analysis of diagnostic studies using exact likelihood

Publication bias (PB) poses a significant threat to meta-analysis, as studies yielding notable results are more likely to be published in scientific journals. Sensitivity analysis provides a flexible method to address PB and to examine the impact of unpublished studies. A selection model based on t-statistics to sensitivity analysis is proposed by Copas. This t-statistics selection model is interpretable and enables the modeling of biased publication sampling across studies, as indicated by the asymmetry in the funnel-plot. In meta-analysis of diagnostic studies, the summary receiver operating characteristic curve is an essential tool for synthesizing the bivariate outcomes of sensitivity and specificity reported by individual studies. Previous studies address PB upon the bivariate normal model but these methods rely on the normal approximation for the empirical logit-transformed sensitivity and specificity, which is not suitable for sparse data scenarios. Compared to the bivariate normal model, the bivariate binomial model which replaces the normal approximation in the within-study model with the exact within-study model has better finite sample properties. In this study, we applied the Copas t-statistics selection model to the meta-analysis of diagnostic studies using the bivariate binomial model. To our knowledge, this is the first study to apply the Copas t-statistics selection model to the bivariate binomial model. We have evaluated our proposed method through several real-world meta-analyses of diagnostic studies and simulation studies.

stat.AP↗

A Practical Analysis Procedure on Generalizing Comparative Effectiveness in the Randomized Clinical Trial to the Real-world Trialeligible Population

When evaluating the effectiveness of a drug, a Randomized Controlled Trial (RCT) is often considered the gold standard due to its perfect randomization. While RCT assures strong internal validity, its restricted external validity poses challenges in extending treatment effects to the broader real-world population due to possible heterogeneity in covariates. In this paper, we introduce a procedure to generalize the RCT findings to the real-world trial-eligible population based on the adaption of existing statistical methods. We utilized the augmented inversed probability of sampling weighting (AIPSW) estimator for the estimation and omitted variable bias framework to assess the robustness of the estimate against the assumption violation caused by potentially unmeasured confounders. We analyzed an RCT comparing the effectiveness of lowering hypertension between Songling Xuemaikang Capsule (SXC), a traditional Chinese medicine (TCM), and Losartan as an illustration. The generalization results indicated that although SXC is less effective in lowering blood pressure than Losartan on week 2, week 4, and week 6, there is no statistically significant difference among the trial-eligible population at week 8, and the generalization is robust against potential unmeasured confounders.

stat.AP↗

Unveiling LLM Evaluation Focused on Metrics: Challenges and Solutions

Natural Language Processing (NLP) is witnessing a remarkable breakthrough driven by the success of Large Language Models (LLMs). LLMs have gained significant attention across academia and industry for their versatile applications in text generation, question answering, and text summarization. As the landscape of NLP evolves with an increasing number of domain-specific LLMs employing diverse techniques and trained on various corpus, evaluating performance of these models becomes paramount. To quantify the performance, it's crucial to have a comprehensive grasp of existing metrics. Among the evaluation, metrics which quantifying the performance of LLMs play a pivotal role. This paper offers a comprehensive exploration of LLM evaluation from a metrics perspective, providing insights into the selection and interpretation of metrics currently in use. Our main goal is to elucidate their mathematical formulations and statistical interpretations. We shed light on the application of these metrics using recent Biomedical LLMs. Additionally, we offer a succinct comparison of these metrics, aiding researchers in selecting appropriate metrics for diverse tasks. The overarching goal is to furnish researchers with a pragmatic guide for effective LLM evaluation and metric selection, thereby advancing the understanding and application of these large language models.

cs.CL↗

Direct and Indirect Treatment Effects in the Presence of Semi-Competing Risks

Semi-competing risks refer to the phenomenon that the terminal event (such as death) can censor the non-terminal event (such as disease progression) but not vice versa. The treatment effect on the terminal event can be delivered either directly following the treatment or indirectly through the non-terminal event. We consider two strategies to decompose the total effect into a direct effect and an indirect effect under the framework of mediation analysis in completely randomized experiments by adjusting the prevalence and hazard of non-terminal events, respectively. They require slightly different assumptions on cross-world quantities to achieve identifiability. We establish asymptotic properties for the estimated counterfactual cumulative incidences and decomposed treatment effects. We illustrate the subtle difference between these two decompositions through simulation studies and two real-data applications.

stat.ME↗