SearcharxivSearch

arXiv subjects

Alex Ocampo

Publications and source records attributed to Alex Ocampo.

9 recordsLinked to original sources

Pseudo-value Based Mean Cumulative Count Regression

The mean cumulative function (MCF) summarizes how events accumulate over time for a recurrent or multi-component endpoint. The MCF, and its integral over a given time horizon, the area under the MCF (AUMCF), provide interpretable summaries of recurrent-event burden in the presence of right-censoring and terminal events. Existing approaches for these estimands have focused primarily on nonparametric treatment comparisons, covariate-adjusted augmentation, and linearized test statistics. Herein, we propose a pseudo-value-based regression approach for estimating covariate effects on the MCF and AUMCF at a fixed truncation time. The proposed method uses influence-function-based pseudo-values as regression outcomes, allowing estimation with standard generalized estimating equation machinery and, under an identity link, ordinary least squares. Through simulation studies, we evaluate estimation accuracy, confidence interval coverage, type I error control, and power across a range of recurrent-event settings. We demonstrate the utility of the proposed covariate adjustment procedure through an application to the ORATORIO clinical trial, evaluating the safety and efficacy of ocrelizumab for the treatment of primary progressive multiple sclerosis. Overall, pseudo-value-based regression provides a simple and interpretable framework for modeling covariate effects on cumulative recurrent-event burden over time.

stat.ME

Assessing covariate-adjusted risk differences in small-sample clinical trials

Binary endpoints are common in clinical trials and conditional odds ratios have traditionally been used to assess treatment effects. However, the interpretation of odds ratios is difficult, they are non-collapsible, and conditional odds-ratios obtained from regression models additionally rely on modeling assumptions in order to be a relevant overall summary measure for the trial. As an alternative, risk differences have gained increasing prominence as a more interpretable, clinically meaningful and assumption-lean measure of treatment effects. This shift has also been motivated by new regulatory guidance, which emphasizes the relevance of marginal estimands and encourages covariate adjustment. Yet, covariate-adjusted inference for risk differences, particularly in smaller samples, has methodological subtleties and lacks well-established best practices. We conduct a simulation study comparing methods for estimating and testing risk differences in small-sample (N$\,\leq\,$150) randomized clinical trials with prognostic categorical baseline covariates, focusing on exact unconditional tests, Mantel-Haenszel methods, and $g$-computation (standardization) approaches. We find that several $g$-computation approaches exhibit inflated Type I error in very small samples when standard Wald-type inference is applied, whereas robust or penalized variants improve error control at the expense of power. Classical methods such as the Mantel-Haenszel and Suissa-Shuster tests remain robust but may forgo efficiency gains from covariate adjustment. Overall, our results suggest that misalignment between estimand and variance estimation may contribute to the Type I error inflation, beyond the impact of small sample size alone. Based on these results, we provide practical recommendations to guide method selection that align the estimand, variance estimation, and inferential target.

stat.ME

Improving Variance Estimation for Covariate Adjustment with Binary Outcomes

Covariate adjustment is a general method for improving precision when estimating treatment effects in randomized trials and is recommended by the FDA in its 2023 guidance when baseline variables are prognostic for the primary outcome. We focus on a method highlighted in that guidance called ``standardization" (or ``g-computation") for estimating the marginal treatment effect. We address the question of how to reliably estimate variance for binary outcomes when marginal outcome probabilities are close to 0 or 1. We propose an influence function-based leave-one-out cross-validated (IF-LOO) variance estimator for the standardized difference-in-means average treatment effect. Through simulation studies, we show that this estimator provides appropriate type-I error control and performs reliably in challenging settings where existing methods can yield inflated type-I error or fail entirely, such as when outcome events are rare or sample sizes are small. In addition to having desirable statistical properties, we derive a closed-form expression for the proposed estimator, enabling straightforward and reliable implementation by study statisticians. The robust finite-sample performance and ease of implementation suggest the IF-LOO variance estimator is a prudent default choice for standardization in clinical trials.

stat.ME

Unified implementation and comparison of Bayesian shrinkage methods for treatment effect estimation in subgroups

Evaluating treatment effect heterogeneity across patient subgroups is a fundamental aspect of clinical trial analysis. These analyses have inherent limitations due to small sample sizes and the substantial number of subgroups investigated. There is a tendency to focus on extreme estimates, which may reflect random variation rather than true effects, potentially leading to spurious clinical conclusions. Statisticians in regulatory agencies and pharmaceutical companies have begun considering shrinkage methods grounded in Bayesian theory. These methods incorporate priors on treatment effect heterogeneity, which shrink subgroup estimates towards the overall treatment effect. Various shrinkage estimators have been proposed, yet it remains unclear which perform best. This work provides a unified presentation and software implementation of shrinkage methods. It also provides simulation comparisons of one-way and global shrinkage methods for two simulation set-ups. One-way models fit a separate shrinkage model for each subgrouping variable while global models include all subgroup indicators. Both can derive standardized subgroup-specific treatment effects. Across all simulation scenarios, shrinkage methods outperformed the standard subgroup estimator in terms of mean squared error. They were also more efficient in identifying a non-efficacious subgroup. Global shrinkage models tended to have smaller mean squared error and less dependence on hyperprior parameters than one-way models, but also exhibited slightly larger bias and worse frequentist coverage of credible intervals. For both models, hyperprior choices anchored in trial assumptions about the anticipated overall treatment effect size performed well. We conclude that some shrinkage is preferable to none and advocate routine inclusion of shrunken estimates in clinical forest plots to facilitate robust decision-making.

stat.ME

Statistical Methodology Groups in the Pharmaceutical Industry

Research and Development is the largest budget position in the pharmaceutical industry, with clinical trials being a critical, yet costly and time-consuming component to inform decisions. Beyond drug efficacy, the probability of success and efficiency of research and development are highly dependent on the approaches used for designing, analyzing, and interpreting clinical trials. Deep understanding of statistical methodology and quantitative approaches is therefore essential. Consequently, dedicated methodology groups have emerged in mid-size and large pharmaceutical companies and CROs. Their remit is to lead the conception and implementation of innovative quantitative methodologies in order to improve drug development, often by addressing complexities or offering more efficient designs. To achieve this, they collaborate internally and externally (e.g., with academics, regulators) to identify common challenges and tear down silos in order to invest in methods with the highest impact on efficiency and value to the portfolio. Given the immense financial stakes of drug development -- where delays carry massive implications -- these groups represent a critical strategic investment. However, to realize this business impact, statistical innovations must be rigorously validated and seamlessly integrated. This manuscript explores the setup, remit, and value of dedicated methodology groups, alongside the critical organizational considerations and success factors required to maximize their impact on the speed, efficiency, and probability of success.

stat.OT

Revealing the Truth: Calculating True Values in Causal Inference Simulation Studies via Gaussian Quadrature

Simulation studies are used to understand the properties of statistical methods. A key luxury in many simulation studies is knowledge of the true value (i.e. the estimand) being targeted. With this oracle knowledge in-hand, the researcher conducting the simulation study can assess across repeated realizations of the data how well a given method recovers the truth. In causal inference simulation studies, the truth is rarely a simple parameter of the statistical model chosen to generate the data. Instead, the estimand is often an average treatment effect, marginalized over the distribution of confounders and/or mediators. Luckily, these variables are often generated from common distributions such as the normal, uniform, exponential, or gamma. For all these distributions, Gaussian quadratures provide efficient and accurate calculation for integrands with integral kernels that stem from known probability density functions. We demonstrate through four applications how to use Gaussian quadrature to accurately and efficiently compute the true causal estimand. We also compare the pros and cons of Gauss-Hermite quadrature to Monte Carlo integration approaches, which we use as benchmarks. Overall, we demonstrate that the Gaussian quadrature is an accurate tool with negligible computation time, yet is underused for calculating the true causal estimands in simulation studies.

stat.ME

Simplifying Causal Mediation Analysis for Time-to-Event Outcomes using Pseudo-Values

Mediation analysis for survival outcomes is challenging. Most existing methods quantify the treatment effect using the hazard ratio (HR) and attempt to decompose the HR into the direct effect of treatment plus an indirect, or mediated, effect. However, the HR is not expressible as an expectation, which complicates this decomposition, both in terms of estimation and interpretation. Here, we present an alternative approach which leverages pseudo-values to simplify estimation and inference. Pseudo-values take censoring into account during their construction, and once derived, can be modeled in the same way as any continuous outcome. Thus, pseudo-values enable mediation analysis for a survival outcome to fit seamlessly into standard mediation software (e.g. CMAverse in R). Pseudo-values are easy to calculate via a leave-one-observation-out procedure (i.e. jackknifing) and the calculation can be accelerated when the influence function of the estimator is known. Mediation analysis for causal effects defined by survival probabilities, restricted mean survival time, and cumulative incidence functions - in the presence of competing risks - can all be performed within this framework. Extensive simulation studies demonstrate that the method is unbiased across 324 scenarios/estimands and controls the type-I error at the nominal level under the null of no mediation. We illustrate the approach using data from the PARADIGMS clinical trial for the treatment of pediatric multiple sclerosis using fingolimod. In particular, we evaluate whether an imaging biomarker lies on the causal path between treatment and time-to-relapse, which aids in justifying this biomarker as a surrogate outcome. Our approach greatly simplifies mediation analysis for survival data and provides a decomposition of the total effect that is both intuitive and interpretable.

stat.ME

Single-World Intervention Graphs for Defining, Identifying, and Communicating Estimands in Clinical Trials

Confusion often arises when attempting to articulate target estimand(s) of a clinical trial in plain language. We aim to rectify this confusion by using a type of causal graph called the Single-World Intervention Graph (SWIG) to provide a visual representation of the estimand that can be effectively communicated to interdisciplinary stakeholders. These graphs not only display estimands, but also illustrate the assumptions under which a causal estimand is identifiable by presenting the graphical relationships between the treatment, intercurrent events, and clinical outcomes. To demonstrate its usefulness in pharmaceutical research, we present examples of SWIGs for various intercurrent event strategies specified in the ICH E9(R1) addendum, as well as an example from a real-world clinical trial for chronic pain. Latex code to generate all the SWIGs shown is this paper is made available. We advocate clinical trialists adopt the use of SWIGs in their estimand discussions during the planning stages of their studies.

stat.ME

Identifying Treatment Effects using Trimmed Means when Data are Missing Not at Random

Patients often discontinue treatment in a clinical trial because their health condition is not improving. Consequently, the patients still in the study at the end of the trial have better health outcomes on average than the initial patient population would have had if every patient had completed the trial. If we only analyze the patients who complete the trial, then this missing data problem biases the estimator of a medication's efficacy because study outcomes are missing not at random (MNAR). One way to overcome this problem - the trimmed means approach for missing data - sets missing values as slightly worse than the worst observed outcome and then trims away a fraction of the distribution from each treatment arm before calculating differences in treatment efficacy (Permutt 2017, Pharmaceutical statistics 16.1:20-28). In this paper we derive sufficient and necessary conditions for when this approach can identify the average population treatment effect in the presence of MNAR data. Numerical studies show the trimmed means approach's ability to effectively estimate treatment efficacy when data are MNAR and missingness is strongly associated with an unfavorable outcome, but trimmed means fail when data are missing at random (MAR) when the better approach would be to multiply impute the missing values. If the reasons for discontinuation in a clinical trial are known analysts can improve estimates with a combination of multiple imputation (MI) and the trimmed means approach when the assumptions of each missing data mechanism hold. When the assumptions are justifiable, using trimmed means can help identify treatment effects notwithstanding MNAR data.

stat.ME