SearcharxivSearch

arXiv subjects

Guoqing Diao

Publications and source records attributed to Guoqing Diao.

11 recordsLinked to original sources

A semiparametric approach for the estimation of covariate-adjusted area under the receiver operating characteristic curve

Receiver operating characteristic (ROC) and the area under the ROC curve (AUC) are widely used to evaluate the discriminative ability of biomarkers. In many clinical settings, however, diagnostic accuracy varies substantially across patient characteristics, and failure to account for such heterogeneity can lead to misleading conclusions. We propose a new semiparametric framework based on generalized additive models to estimate covariate-specific and covariate-adjusted AUC while allowing for nonlinear and interaction effects of covariates on biomarker performance. Our method accommodates both binary and multicategory disease status. We also establish the asymptotic properties of the proposed estimators. Simulations demonstrate favorable finite-sample performance. We illustrate the method using data from the Alzheimer's Disease Neuroimaging Initiative, where substantial heterogeneity in biomarker discrimination across covariates is observed.

stat.ME

A Unified Causal Inference Framework for the Desirability of Outcome Ranking Paradigm in Benefit-Risk Evaluation

We developed a unified covariate-adjusted causal inference framework for estimating the desirability of outcome ranking (DOOR) probability for benefit-risk evaluation in randomized trials and observational studies. The framework expresses the DOOR probability as a bilinear functional of the marginal ordinal outcome distributions under the two treatment strategies, estimates conditional ordinal distributions through sequential risk-set hazards, and derives the efficient influence function (EIF) of the DOOR probability. The point-estimation simulations compared G-computation, normalized inverse probability weighting (IPW), augmented IPW (AIPW), and targeted maximum likelihood estimation (TMLE), with nuisance functions estimated using generalized linear models or Super Learner (SL). TMLE-SL showed the strongest and most consistent point-estimation performance, with AIPW-SL ranking second. EIF-based inference was then evaluated for AIPW-SL and TMLE-SL, with and without cross-fitting, across settings varying in overlap, treatment-effect heterogeneity, and treatment allocation. CVTMLE-SL showed the strongest overall performance across DOOR-scale bias, recovery of the underlying ordinal distributions, standard-error accuracy, and confidence-interval coverage. We illustrate the methodology using data from the multidrug-resistant organism network of the Antibacterial Resistance Leadership Group.

stat.ML

On Cluster Randomized Trials with the Desirability of Outcome Ranking (DOOR) Endpoints

Cluster randomized trials are widely used when individual randomization is logistically infeasible or when correlations between observations cannot be ignored, especially in fields such as ophthalmology, infectious disease, vaccine research, and sociology. The desirability of outcome ranking (DOOR) framework evaluates patient-centric benefit-risk using an ordinal outcome and a Wilcoxon-Mann-Whitney statistic-based approach to compare outcome distributions between interventions. We propose a suite of new methods to extend DOOR to cluster trials based on properties of U-statistics and influence functions to estimate within-cluster and between-cluster treatment effects. These approaches can be applied in different scenarios, including mixtures of clusters with two treatment groups and clusters with only one group, and both small and large numbers of clusters. Simulations demonstrate that the proposed methods perform well under various scenarios regarding the number of clusters and cluster sizes. As an illustration, we apply the proposed methods to a cluster randomized crossover trial comparing delayed cord clamping and umbilical cord milking for newborns.

stat.ME

A likelihood approach to proper analysis of secondary outcomes in matched case-control studies

Matched case-control studies are commonly employed in epidemiological research for their convenience and efficiency. Analysis of secondary outcomes can yield valuable insights into biological pathways and help identify genetic variants of importance. Naive analysis using standard statistical methods, such as least-squares regression for quantitative traits, can be misleading because they fail to account for unequal sampling induced by the case-control design and matching. In this paper, we propose novel statistical methods that appropriately reflect the study design and sampling scheme in the analysis of secondary outcome data. The new methods provide consistent estimation and accurate coverage probabilities for the confidence interval estimators. We demonstrate the advantages of the new methods through simulation studies and a real application with diabetes patients. R code implementing the proposed methods is publicly available.

stat.ME

Doubly Robust Estimation of Desirability of Outcome Ranking (DOOR) Probability with Application to MDRO Studies

In observational studies, adjusting for confounders is required if a treatment comparison is planned. A crude comparison of the primary endpoint without covariate adjustment will suffer from biases, and the addition of regression models could improve precision by incorporating imbalanced covariates and thus help make correct inference. Desirability of outcome ranking (DOOR) is a patient-centric benefit-risk evaluation methodology designed for randomized clinical trials. Still, robust covariate adjustment methods could further expand the compatibility of this method in observational studies. In DOOR analysis, each participant's outcome is ranked based on pre-specified clinical criteria, where the most desirable rank represents a good outcome with no side effects and the least desirable rank is the worst possible clinical outcome. We develop a causal framework for estimating the population-level DOOR probability, via the inverse probability of treatment weighting method, G-Computation method, and a Doubly Robust method that combines both. The performance of the proposed methodologies is examined through simulations. We also perform a causal analysis of the Multi-Drug Resistant Organism (MDRO) network within the Antibacterial Resistant Leadership Group (ARLG), comparing the benefit:risk between Mono-drug therapy and Combination-drug therapy.

stat.ME

Desirability of outcome ranking (DOOR) analysis for multivariate survival outcomes with application to ACTT-1 trial

Desirability Of Outcome Ranking (DOOR) methodology accounts for problems that conventional benefit:risk analyses in clinical trials ignore, such as competing risks and the trade-off relationship between efficacy and toxicity. DOOR levels can be considered as a multi-state process in nature, as event-free survival, and survival with side effects are not equivalent and the overall patient trajectory requires recognition. In monotone settings where patients' conditions can only decline, we can record event times for each transition from one level of the DOOR to another, and construct Kaplan-Meier curves displaying transition times. While traditional survival analysis methods such as the Cox model require assumptions like proportional hazards and suffer from the challenge of interpreting a hazard ratio, Restricted Mean Survival Time (RMST) offers an alternative with greater intuitiveness. Therefore, we propose a combination of the two domains to develop estimation and inferential procedures that could benefit from the advantages of both DOOR and RMST. Particularly, the area under each survival curve restricted to a time point, or the RMST, has clear clinical meanings, from expected event-free survival time, expected survival time with at most one of the events, to expected lifetime before death. We show that the nonparametric estimator of the RMSTs asymptotically follows a multivariate Gaussian process through the martingale theory and functional delta method. There are alternative approaches to hypothesis testing that recognize when patients transition into worse states. We evaluate our proposed method with data simulated under a multistate model. We consider various scenarios, including when the null hypothesis is true, when the treatment difference exists only in certain DOOR levels, and small-sample studies. We also present a real-world example with ACTT-1.

stat.ME

Fast Algorithm for Calculating Probability of Chess Winning Streaks

Motivated by the controversy in the chess community, where Hikaru Nakamura, a renowned grandmaster, has posted multiple impressive winning streaks over the years on the online platform chess.com, we derive the probabilities of various types of streaks in online chess and/or other sports. Specifically, given the winning/drawing/losing probabilities of individual games, we derive the probabilities of "pure" winning streaks, non-losing streaks, and "in-between" streaks involving at most one draw over the course of games played in a period. The performance of the developed algorithms is examined through numerical studies.

math.HO

A General Framework for Regression with Mismatched Data Based on Mixture Modeling

Data sets obtained from linking multiple files are frequently affected by mismatch error, as a result of non-unique or noisy identifiers used during record linkage. Accounting for such mismatch error in downstream analysis performed on the linked file is critical to ensure valid statistical inference. In this paper, we present a general framework to enable valid post-linkage inference in the challenging secondary analysis setting in which only the linked file is given. The proposed framework covers a wide selection of statistical models and can flexibly incorporate additional information about the underlying record linkage process. Specifically, we propose a mixture model for pairs of linked records whose two components reflect distributions conditional on match status, i.e., correct match or mismatch. Regarding inference, we develop a method based on composite likelihood and the EM algorithm as well as an extension towards a fully Bayesian approach. Extensive simulations and several case studies involving contemporary record linkage applications corroborate the effectiveness of our framework.

stat.ME

A Pseudo-Likelihood Approach to Linear Regression with Partially Shuffled Data

Recently, there has been significant interest in linear regression in the situation where predictors and responses are not observed in matching pairs corresponding to the same statistical unit as a consequence of separate data collection and uncertainty in data integration. Mismatched pairs can considerably impact the model fit and disrupt the estimation of regression parameters. In this paper, we present a method to adjust for such mismatches under ``partial shuffling" in which a sufficiently large fraction of (predictors, response)-pairs are observed in their correct correspondence. The proposed approach is based on a pseudo-likelihood in which each term takes the form of a two-component mixture density. Expectation-Maximization schemes are proposed for optimization, which (i) scale favorably in the number of samples, and (ii) achieve excellent statistical performance relative to an oracle that has access to the correct pairings as certified by simulations and case studies. In particular, the proposed approach can tolerate considerably larger fraction of mismatches than existing approaches, and enables estimation of the noise level as well as the fraction of mismatches. Inference for the resulting estimator (standard errors, confidence intervals) can be based on established theory for composite likelihood estimation. Along the way, we also propose a statistical test for the presence of mismatches and establish its consistency under suitable conditions.

stat.ME

Rare event simulation for processes generated via stochastic fixed point equations

In a number of applications, particularly in financial and actuarial mathematics, it is of interest to characterize the tail distribution of a random variable $V$ satisfying the distributional equation $V\stackrel{\mathcal{D}}{=}f(V)$, where $f(v)=A\max\{v,D\}+B$ for $(A,B,D)\in(0,\infty)\times {\mathbb{R}}^2$. This paper is concerned with computational methods for evaluating these tail probabilities. We introduce a novel importance sampling algorithm, involving an exponential shift over a random time interval, for estimating these rare event probabilities. We prove that the proposed estimator is: (i) consistent, (ii) strongly efficient and (iii) optimal within a wide class of dynamic importance sampling estimators. Moreover, using extensions of ideas from nonlinear renewal theory, we provide a precise description of the running time of the algorithm. To establish these results, we develop new techniques concerning the convergence of moments of stopped perpetuity sequences, and the first entrance and last exit times of associated Markov chains on $\mathbb{R}$. We illustrate our methods with a variety of numerical examples which demonstrate the ease and scope of the implementation.

math.PR

Efficient Semiparametric Estimation of Short-term and Long-term Hazard Ratios with Right-Censored Data

The proportional hazards assumption in the commonly used Cox model for censored failure time data is often violated in scientific studies. Yang and Prentice (2005) proposed a novel semiparametric two-sample model that includes the proportional hazards model and the proportional odds model as sub-models, and accommodates crossing survival curves. The model leaves the baseline hazard unspecified and the two model parameters can be interpreted as the short-term and long-term hazard ratios. Inference procedures were developed based on a pseudo score approach. Although extension to accommodate covariates was mentioned, no formal procedures have been provided or proved. Furthermore, the pseudo score approach may not be asymptotically efficient. We study the extension of the short-term and long-term hazard ratio model of Yang and Prentice (2005) to accommodate potentially time-dependent covariates. We develop efficient likelihood-based estimation and inference procedures. The nonparametric maximum likelihood estimators are shown to be consistent, asymptotically normal, and asymptotically efficient. Extensive simulation studies demonstrate that the proposed methods perform well in practical settings. The proposed method captured the phenomenon of crossing hazards in a cancer clinical trial and identified a genetic marker with significant long-term effect missed by using the proportional hazards model on age-at-onset of alcoholism in a genetic study.

stat.ME