SearcharxivSearch

arXiv subjects

Bryan E. Shepherd

Publications and source records attributed to Bryan E. Shepherd.

At least 19 recordsLinked to original sources

Authentic Multinational Federated Time-to-Event Analyses Among People with HIV in Latin America

Multinational HIV cohort studies face regulatory barriers to cross-border sharing of individual participant data, limiting centralized pooled analyses. Federated statistical methods, which exchange only aggregated information, offer a privacy-preserving alternative but have rarely been examined in real-world distributed environments for HIV research. Here, we evaluate the feasibility and analytical performance of a communication-efficient federated framework within the Caribbean, Central, and South America Network for HIV epidemiology. Virologic failure and major regimen change after antiretroviral therapy initiation were analyzed as separate outcomes; for each, we estimated cumulative incidence functions (CIFs) and fit stratified cause-specific Cox proportional hazards models via a surrogate likelihood-based federated implementation. Each site imputed missing data, conducted local analysis, and shared only summary statistics according to a coordinated computation protocol. The federated approach exactly reproduced centralized CIFs and closely approximated centralized Cox regression estimates, outperforming conventional meta-analysis for both outcomes. These findings demonstrate that authentic federated analysis is feasible for multinational HIV research and can yield results closely aligned with centralized analysis while preserving data privacy. Post-hoc feedback from local analysts, however, indicated that broader adoption will require managing the logistical and coordination overhead and ensuring harmonized data collection and quality control across sites.

stat.AP

Doubly Robust Estimators of Quantile Treatment Effects With Semiparametric Cumulative Probability Models

The causal inference literature has traditionally focused on estimating the mean of the potential outcome, whereas evaluating how a treatment affects the entire outcome distribution can provide additional information in biomedical research. Quantile treatment effect (QTE) captures such distributional differences, particularly when outcomes are skewed. However, existing approaches for estimating QTE make distributional assumptions about the outcome and are thus sensitive to model misspecification. Motivated by an HIV study with skewed outcomes, one of which is subject to detection limits, we propose a doubly robust framework for estimating QTE based on the cumulative probability model (CPM), which is a rank-based, semiparametric linear transformation model. We develop two CPM-based estimation strategies: (1) an inverse-cumulative distribution function (CDF) approach that first estimates the marginal CDF of potential outcomes using the efficient influence function (EIF) and then obtains marginal quantiles via weighted quantile interpolation by inverting the distribution, and (2) a direct approach that solves the EIF of potential marginal quantiles. The proposed estimators are doubly robust and asymptotically normal. We further extend the framework to probability treatment effects (PTEs) and their conditional counterparts. For statistical inference, we investigate several variance estimation procedures, including EIF-based estimators, sandwich estimators, and the nonparametric bootstrap. Simulation studies illustrate that the empirical sandwich estimator and the nonparametric bootstrap provide doubly robust variance estimation with stable finite-sample performance under nuisance model misspecification. The proposed methods are evaluated through extensive Monte Carlo simulations and illustrated using an HIV data application.

stat.ME

Methods to address measurement error in both Outcome and Covariates

Biomedical research is increasingly relying on readily available routine data, such as electronic health records. Routinely collected data, as well as datasets from large cohorts, are often prone to measurement error which, if not addressed in analyses, can bias study results and ultimately mislead clinical decision-making and potentially harm patients. For this setting, methods that address errors in the outcome and multiple covariates are needed. In this tutorial, we will review available methods to address for errors in both outcomes and covariates. We will illustrate methods with use of a running example in order to compare the methods directly. Both the data and analytic code are provided for the user so that they may easily reproduce results in each example. We conclude the tutorial with a discussion of the different approaches and highlight areas of future work needed for this setting.

stat.AP

Addressing errors in multiple variables using generalized raking and cumulative probability models

Routinely collected data, such as electronic health record (EHR) data, are frequently used for biomedical research, but these data are prone to errors, which can bias study findings. Validating data in subsamples of records can reduce bias, and the efficiency of estimates can be improved by incorporating in analyses both the error-prone data available on the entire cohort and the validated data available on the subsample. One approach to incorporate both data sources is with generalized raking, which calibrates validation sampling weights using error-prone data from the entire cohort. Motivated by an EHR study of maternal weight gain during pregnancy with a validation subsample, we develop and illustrate generalized raking techniques for cumulative probability models (CPMs). CPMs are robust, rank-based and semiparametric models for continuous, ordinal, or mixed type outcome data. We develop efficient generalized raking estimators for CPMs, evaluate their performance relative to competing methods, and demonstrate the utility and strengths of generalized raking with CPMs in a study that examines factors associated with weight gain during pregnancy.

stat.ME

Generalized raking and stabilized weights for regression modeling in two-phase samples

In regression models fitted to data from complex survey designs, sampling weights often incorporate non-essential variation, inflating variance estimates. Stabilized weights mitigate this issue by adjusting sampling weights to account for variation explained by covariates. In the context of two-phase sampling, we evaluate the performance of optimal stabilized weights and propose combining the stabilized weight estimator with generalized raking, a class of efficient design-based estimators. This combination improves efficiency by reducing unnecessary weight variation and leveraging information from auxiliary variables. We show this combination can be implemented using the standard statistical package that handles two-phase samples and generalized raking. Simulation studies demonstrate that the proposed estimator enhances precision under realistic two-phase designs, though efficiency gains may be limited in highly informative designs. The developed methods were applied to a large multinational two-phase study of Kaposi sarcoma among people living with HIV.

stat.ME

Improving optimal subsampling through stratification

Recent works have proposed optimal subsampling algorithms to improve computational efficiency in large datasets and to design validation studies in the presence of measurement error. Existing approaches generally fall into two categories: (i) designs that optimize individualized sampling rules, where unit-specific probabilities are assigned and applied independently, and (ii) designs based on stratified sampling with simple random sampling within strata. Focusing on the logistic regression setting, we derive the asymptotic variances of estimators under both approaches and compare them numerically through extensive simulations and an application to data from the Vanderbilt Comprehensive Care Clinic cohort. Our results reinforce that stratified sampling is not merely an approximation to individualized sampling, showing instead that optimal stratified designs are often more efficient than optimal individualized designs through their elimination of between-stratum contributions to variance. These findings suggest that optimizing over the class of individualized sampling rules overlooks highly efficient sampling designs and highlight the often underappreciated advantages of stratified sampling.

stat.ME

Optimal two-phase sampling designs for generalized raking estimators with multiple parameters of interest

Large observational datasets, including those derived from electronic health records, are a valuable resource for medical research but are often affected by missingness, measurement error, and misclassification. Two-phase sampling with generalized raking (GR) estimation is an efficient and robust approach to statistical inference in such settings. In this approach, variables that are unavailable or measured with error in a large phase 1 cohort are obtained with higher-quality measurements in a phase 2 subsample. Previous research has studied optimal phase 2 sampling designs for inverse probability weighted (IPW) estimators in non-adaptive, multi-parameter settings, and for GR estimators in single-parameter settings. In this work, we extend these results by deriving optimal adaptive, multiwave sampling designs for IPW and GR estimators when multiple parameters are of interest. We propose several practical allocation strategies and evaluate their performance through extensive simulations and a data example from the Vanderbilt Comprehensive Care Clinic HIV Study. Our results show that independently optimizing allocation for each parameter improves efficiency over traditional case-control sampling. We also derive an integer-valued, A-optimal allocation method that typically outperforms independent optimization. Notably, we find that optimal designs for GR can differ substantially from those for IPW, and that this distinction can meaningfully affect estimator efficiency in the multiple-parameter setting. These findings offer practical guidance for future two-phase studies involving incomplete or error-prone data.

stat.ME

Flexible and Efficient Estimation of Causal Effects with Error-Prone Exposures: A Control Variates Approach for Measurement Error

Exposure measurement error is a ubiquitous but often overlooked challenge in causal inference with observational data. Existing methods accounting for exposure measurement error largely rely on restrictive parametric assumptions, while emerging data-adaptive estimation approaches allow for less restrictive assumptions but at the cost of flexibility, as they are typically tailored towards rigidly-defined statistical quantities. There remains a critical need for assumption-lean estimation methods that are both flexible and possess desirable theoretical properties across a variety of study designs. In this paper, we introduce a general framework for estimation of causal quantities in the presence of exposure measurement error, adapted from the control variates approach of Yang and Ding (2019). Our method can be implemented in various two-phase sampling study designs, where one obtains gold-standard exposure measurements for a small subset of the full study sample, called the validation data. The control variates framework leverages both the error-prone and error-free exposure measurements by augmenting an initial consistent estimator from the validation data with a variance reduction term formed from the full data. We show that our method inherits double-robustness properties under standard causal assumptions. Simulation studies show that our approach performs favorably compared to leading methods under various two-phase sampling schemes. We illustrate our method with observational electronic health record data on HIV outcomes from the Vanderbilt Comprehensive Care Clinic.

stat.ME

Efficient Estimation of Causal Effects Under Two-Phase Sampling with Error-Prone Outcome and Treatment Measurements

Measurement error is a common challenge for causal inference studies using electronic health record (EHR) data, where clinical outcomes and treatments are frequently mismeasured. Researchers often address measurement error by conducting manual chart reviews to validate measurements in a subset of the full EHR data -- a form of two-phase sampling. To improve efficiency, phase-two samples are often collected in a biased manner dependent on the patients' initial, error-prone measurements. In this work, motivated by our aim of performing causal inference with error-prone outcome and treatment measurements under two-phase sampling, we develop solutions applicable to both this specific problem and the broader problem of causal inference with two-phase samples. For our specific measurement error problem, we construct two asymptotically equivalent doubly-robust estimators of the average treatment effect and demonstrate how these estimators arise from two previously disconnected approaches to constructing efficient estimators in general two-phase sampling settings. We document various sources of instability affecting estimators from each approach and propose modifications that can considerably improve finite sample performance in any two-phase sampling context. We demonstrate the utility of our proposed methods through simulation studies and an illustrative example assessing effects of antiretroviral therapy on occurrence of AIDS-defining events in patients with HIV from the Vanderbilt Comprehensive Care Clinic.

stat.ME

Estimating treatment effects with a unified semi-parametric difference-in-differences approach

Difference-in-differences (DID) approaches are widely used for estimating causal effects with observational data before and after an intervention. DID traditionally estimates the average treatment effect among the treated after making a parallel trends assumption on the means of the outcome. With skewed outcomes, a transformation is often needed; however, the transformation may be difficult to choose, results may be sensitive to the choice, and parallel trends assumptions are made on the transformed scale. Recent DID methods estimate alternative treatment effects that may be preferable with skewed outcomes. However, each alternative DID estimator requires a different parallel trends assumption. We introduce a new DID method capable of estimating average, quantile, probability, and novel Mann-Whitney treatment effects among the treated with a single unifying parallel trends assumption. The proposed method uses a semi-parametric cumulative probability model (CPM). The CPM is a linear model for a latent variable on covariates, where the latent variable results from an unspecified transformation of the outcome. Our DID approach makes a universal parallel trends assumption on the expectation of the latent variable conditional on covariates. Hence, our method avoids specifying outcome transformations and does not require separate assumptions for each estimand. We introduce the method; describe identification, estimation, and inference; conduct simulations evaluating its performance; and apply it to assess the impact of Medicaid expansion on CD4 count among people with HIV.

stat.ME

Double machine learning to estimate the effects of multiple treatments and their interactions

Causal inference literature has extensively focused on binary treatments, with relatively fewer methods developed for multi-valued treatments. In particular, methods for multiple simultaneously assigned treatments remain understudied despite their practical importance. This paper introduces two settings: (1) estimating the effects of multiple treatments of different types (binary, categorical, and continuous) and the effects of treatment interactions, and (2) estimating the average treatment effect across categories of multi-valued regimens. To obtain robust estimates for both settings, we propose a class of methods based on the Double Machine Learning (DML) framework. Our methods are well-suited for complex settings of multiple treatments/regimens, using machine learning to model confounding relationships while overcoming regularization and overfitting biases through Neyman orthogonality and cross-fitting. To our knowledge, this work is the first to apply machine learning for robust estimation of interaction effects in the presence of multiple treatments. We further establish the asymptotic distribution of our estimators and derive variance estimators for statistical inference. Extensive simulations demonstrate the performance of our methods. Finally, we apply the methods to study the effect of three treatments on HIV-associated kidney disease in an adult HIV cohort of 2455 participants in Nigeria.

stat.ME

Unified and Simple Sample Size Calculations for Individual or Cluster Randomized Trials with Skewed or Ordinal Outcomes

Sample size calculations can be challenging with skewed continuous outcomes in randomized controlled trials (RCTs). Standard t-test-based calculations may require data transformation, which may be difficult before data collection. Calculations based on individual and clustered Wilcoxon rank-sum tests have been proposed as alternatives, but these calculations for clustered data assume no ties in continuous outcomes, and clustered Wilcoxon rank-sum tests perform poorly with heterogeneous cluster sizes. Recent work has shown that continuous outcomes can be robustly analyzed using ordinal cumulative probability models. Analogously, sample size calculations for ordinal outcomes can be a robust design strategy for continuous outcomes. We show that Whitehead sample size calculations for independent ordinal outcomes can naturally extend to continuous outcomes. We extend these calculations to cluster RCTs using a design effect incorporating rank intraclass correlation coefficients. Therefore, we provide a unifying, simple approach for designing individual and cluster RCTs for continuous or ordinal outcomes that makes minimal assumptions on the distribution of the still-to-be-collected outcome. We conduct simulations to evaluate our approach performance and illustrate its application in multiple RCTs: an individual RCT with skewed continuous outcomes, a cluster RCT with skewed continuous outcomes, and a non-inferiority cluster RCT with an irregularly distributed count outcome.

stat.ME

Between- and Within-Cluster Spearman Rank Correlations

Clustered data are common in practice. Clustering arises when subjects are measured repeatedly, or subjects are nested in groups (e.g., households, schools). It is often of interest to evaluate the correlation between two variables with clustered data. There are three commonly used Pearson correlation coefficients (total, between-, and within-cluster), which together provide an enriched perspective of the correlation. However, these Pearson correlation coefficients are sensitive to extreme values and skewed distributions. They also vary with data transformation, which is arbitrary and often difficult to choose, and they are not applicable to ordered categorical data. Current nonparametric correlation measures for clustered data are only for the total correlation. Here we define population parameters for the between- and within-cluster Spearman rank correlations. The definitions are natural extensions of the Pearson between- and within-cluster correlations to the rank scale. We show that the total Spearman rank correlation approximates a linear combination of the between- and within-cluster Spearman rank correlations, where the weights are functions of rank intraclass correlations of the two random variables. We also discuss the equivalence between the within-cluster Spearman rank correlation and the covariate-adjusted partial Spearman rank correlation. Furthermore, we describe estimation and inference for the three Spearman rank correlations, conduct simulations to evaluate the performance of our estimators, and illustrate their use with data from a longitudinal biomarker study and a clustered randomized trial.

stat.ME

Assessing treatment effects in observational data with missing confounders: A comparative study of practical doubly-robust and traditional missing data methods

In pharmacoepidemiology, safety and effectiveness are frequently evaluated using readily available administrative and electronic health records data. In these settings, detailed confounder data are often not available in all data sources and therefore missing on a subset of individuals. Multiple imputation (MI) and inverse-probability weighting (IPW) are go-to analytical methods to handle missing data and are dominant in the biomedical literature. Doubly-robust methods, which are consistent under fewer assumptions, can be more efficient with respect to mean-squared error. We discuss two practical-to-implement doubly-robust estimators, generalized raking and inverse probability-weighted targeted maximum likelihood estimation (TMLE), which are both currently under-utilized in biomedical studies. We compare their performance to IPW and MI in a detailed numerical study for a variety of synthetic data-generating and missingness scenarios, including scenarios with rare outcomes and a high missingness proportion. Further, we consider plasmode simulation studies that emulate the complex data structure of a large electronic health records cohort in order to compare anti-depressant therapies in a rare-outcome setting where a key confounder is prone to more than 50\% missingness. We provide guidance on selecting a missing data analysis approach, based on which methods excelled with respect to the bias-variance trade-off across the different scenarios studied.

stat.ME

Probability-scale residuals for event-time data

The probability-scale residual (PSR) is defined as $E\{sign(y, Y^*)\}$, where $y$ is the observed outcome and $Y^*$ is a random variable from the fitted distribution. The PSR is particularly useful for ordinal and censored outcomes for which fitted values are not available without additional assumptions. Previous work has defined the PSR for continuous, binary, ordinal, right-censored, and current status outcomes; however, development of the PSR has not yet been considered for data subject to general interval censoring. We develop extensions of the PSR, first to mixed-case interval-censored data, and then to data subject to several types of common censoring schemes. We derive the statistical properties of the PSR and show that our more general PSR encompasses several previously defined PSR for continuous and censored outcomes as special cases. The performance of the residual is illustrated in real data from the Caribbean, Central, and South American Network for HIV Epidemiology.

stat.ME

Understanding Difference-in-differences methods to evaluate policy effects with staggered adoption: an application to Medicaid and HIV

While a randomized control trial is considered the gold standard for estimating causal treatment effects, there are many research settings in which randomization is infeasible or unethical. In such cases, researchers rely on analytical methods for observational data to explore causal relationships. Difference-in-differences (DID) is one such method that, most commonly, estimates a difference in some mean outcome in a group before and after the implementation of an intervention or policy and compares this with a control group followed over the same time (i.e., a group that did not implement the intervention or policy). Although DID modeling approaches have been gaining popularity in public health research, the majority of these approaches and their extensions are developed and shared within the economics literature. While extensions of DID modeling approaches may be straightforward to apply to observational data in any field, the complexities and assumptions involved in newer approaches are often misunderstood. In this paper, we focus on recent extensions of the DID method and their relationships to linear models in the setting of staggered treatment adoption over multiple years. We detail the identification and estimation of the average treatment effect among the treated using potential outcomes notation, highlighting the assumptions necessary to produce valid estimates. These concepts are described within the context of Medicaid expansion and retention in care among people living with HIV (PWH) in the United States. While each DID approach is potentially valid, understanding their different assumptions and choosing an appropriate method can have important implications for policy-makers, funders, and public health as a whole.

stat.ME

Rank Intraclass Correlation for Clustered Data

Clustered data are common in biomedical research. Observations in the same cluster are often more similar to each other than to observations from other clusters. The intraclass correlation coefficient (ICC), first introduced by R. A. Fisher, is frequently used to measure this degree of similarity. However, the ICC is sensitive to extreme values and skewed distributions, and depends on the scale of the data. It is also not applicable to ordered categorical data. We define the rank ICC as a natural extension of Fisher's ICC to the rank scale, and describe its corresponding population parameter. The rank ICC is simply interpreted as the rank correlation between a random pair of observations from the same cluster. We also extend the definition when the underlying distribution has more than two hierarchies. We describe estimation and inference procedures, show the asymptotic properties of our estimator, conduct simulations to evaluate its performance, and illustrate our method in three real data examples with skewed data, count data, and three-level ordered categorical data.

stat.ME

Asymptotic Properties for Cumulative Probability Models for Continuous Outcomes

Regression models for continuous outcomes often require a transformation of the outcome, which the user either specify {\it a priori} or estimate from a parametric family. Cumulative probability models (CPMs) nonparametrically estimate the transformation and are thus a flexible analysis approach for continuous outcomes. However, it is difficult to establish asymptotic properties for CPMs due to the potentially unbounded range of the transformation. Here we show asymptotic properties for CPMs when applied to slightly modified data where the outcomes are censored at the ends. We prove uniform consistency of the estimated regression coefficients and the estimated transformation function over the non-censored region, and describe their joint asymptotic distribution. We show with simulations that results from this censored approach and those from the CPM on the original data are similar when a small fraction of data are censored. We reanalyze a dataset of HIV-positive patients with CPMs to illustrate and compare the approaches.

stat.ME