SearcharxivSearch

arXiv subjects

Masahiko Gosho

Publications and source records attributed to Masahiko Gosho.

7 recordsLinked to original sources

Validity of MMRM-based hypothesis testing under missing-not-at-random mechanisms

In randomized clinical trials with longitudinal continuous outcomes, missing-not-at-random (MNAR) missingness often motivates conservative alternatives to mixed models for repeated measures (MMRM). Such caution is important for estimation, but estimation and testing need not require identical assumptions. Moreover, overly conservative primary analyses may reduce power, increase required sample size, and raise trial costs. We investigated the validity of MMRM-based testing under the global null of identical longitudinal outcome distributions across groups. Because valid testing minimally requires treatment-effect estimators to converge to the null under the null hypothesis, we investigated sufficient conditions for this property. We introduced a proportional observation condition requiring ratios of observation probabilities relative to a reference group, conditional on the full outcome vector, to be outcome-independent, and showed that, with arbitrary post-baseline visits and monotone missingness, this condition is sufficient for convergence to the null value. The condition allows observation to depend on unobserved outcomes and permits between-group differences in overall observation probabilities through outcome-independent dropout, making it clinically interpretable while accommodating outcome-dependent MNAR missingness. Synthetic and data-based bootstrap simulations showed negligible bias and empirical test sizes near 0.05, including nonmonotone missingness. Thus, MNAR missingness does not by itself imply that a more conservative primary testing procedure is required. This result does not justify treatment-effect estimation under alternatives, which still requires estimand-based interpretation and sensitivity analyses.

stat.ME

A flexible framework for treatment effect inference in longitudinal clinical studies with skewed outcomes

Longitudinal continuous outcomes in clinical trials are commonly analyzed using mixed models for repeated measures (MMRM) under normality assumptions. However, many clinical outcomes are skewed, making mean-based treatment effects difficult to interpret and potentially reducing statistical efficiency. The Box--Cox MMRM (BCMMRM) approach accommodates skewness by enabling inference on model-based median differences via inverse transformation. However, BCMMRM typically assumes a common transformation parameter across treatment groups and time points. When distributional shapes differ between groups or evolve over time, this assumption may lead to biased treatment effect. Furthermore, when treatment affects not only central tendency but also distributional shape or tail behavior, treatment effects may not be adequately characterized by a single location summary such as the median. We propose the Box--Cox multivariate regression (BCMVR) framework for longitudinal data with skewed outcomes. BCMVR relaxes this restriction by allowing transformation parameters to vary across groups and time points. The framework enables inference based on interpretable summaries, including median differences and a probability-based treatment effect quantifying the probability that a randomly selected patient in one group has a better outcome than one in another group. This measure integrates information over the entire outcome distribution and provides a complementary summary when distributional shapes differ. Simulation studies demonstrate that BCMMRM can produce biased estimates when distributions differ in shape, whereas BCMVR provides nearly unbiased estimation. The probability-based measure achieves a favorable balance between robustness and statistical efficiency. The proposed framework provides a flexible and interpretable approach to treatment effect inference under distributional heterogeneity.

stat.ME

Modification and extension of the Bayesian clinical trial design using external data for single-arm and hybrid-controlled trials

Limited patient availability complicates sample size determination in pediatric clinical trials. Although Bayesian methods incorporating external data offer a solution, rigorously controlling the type I error rate remains difficult. Psioda and Ibrahim (2019) proposed a simulation-based framework as a practical solution. However, although their framework was designed to relax the type I error control, this relaxation fails when the external data exhibit a large treatment effect, making it difficult to design clinical trials that incorporate external data. Furthermore, restricting the support of sampling priors can cause trial outcomes to fall outside of this support, leading to lower power. Additionally, their analytic prior formulation may induce bias, and their method is not applicable to hybrid-controlled trials involving two-group comparisons. Thus, we propose modifications to both the sampling and analytic prior specifications and extend the framework to hybrid-controlled trials. We redefine the null sampling prior as a normal distribution centered at the null boundary, ensuring a Bayesian type I error evaluation. For the analytic prior, we employ a weakly informative prior for the second component of a robust mixture prior to mitigate bias under prior-data conflict. Furthermore, we extend this methodology to hybrid-controlled trials. Simulation studies and a pediatric case study of cutaneous lupus erythematosus demonstrate that our method substantially reduces the required sample size compared with both frequentist and original Bayesian methods, while maintaining the target operating characteristics and controlling estimation bias under prior-data conflict. This framework provides a reliable and efficient approach for designing clinical trials that incorporate external information.

stat.ME

Posterior Quantification of Borrowing from Multiple Historical Control Data in Bayesian Dynamic Borrowing Methods: A Scoping Review

Bayesian dynamic borrowing methods incorporate historical control data into current clinical trial analyses while allowing the degree of borrowing to depend on the compatibility between historical and current data. Although many methods have been proposed, the degree of borrowing is often difficult to interpret, especially when multiple historical control sources are available. This scoping review focuses on posterior quantification of borrowing from multiple historical controls. We discuss overall borrowing summaries based on effective historical sample size, together with method-specific source-level summaries of borrowing, information contribution, or compatibility arising from power priors, unit information priors, multisource exchangeability models, Dirichlet process mixture models, and potential bias models. We distinguish posterior borrowing measures from quantities describing prior information allocation or source-specific conflict. Two case studies, one with a binary endpoint and one with a continuous endpoint, illustrate that methods with broadly similar posterior treatment effect estimates may differ in both the overall amount and source-specific pattern of borrowing. These examples show that large overall borrowing may reflect selective borrowing from compatible historical sources rather than uniform borrowing from all sources. We recommend reporting treatment effect estimates together with overall and source-specific borrowing summaries, when available, to improve transparency in posterior inference.

stat.ME

Nonparametric Bayesian approach for dynamic borrowing of historical control data

When incorporating historical control data into the analysis of current randomized controlled trial data, it is critical to account for differences between the datasets. When the cause of the difference is an unmeasured factor and adjustment for observed covariates only is insufficient, it is desirable to use a dynamic borrowing method that reduces the impact of heterogeneous historical controls. We propose a nonparametric Bayesian approach for borrowing historical controls that are homogeneous with the current control. Additionally, to emphasize the resolution of conflicts between the historical controls and current control, we introduce a method based on the dependent Dirichlet process mixture. The proposed methods can be implemented using the same procedure, regardless of whether the outcome data comprise aggregated study-level data or individual participant data. We also develop a novel index of similarity between the historical and current control data, based on the posterior distribution of the parameter of interest. We conduct a simulation study and analyze clinical trial examples to evaluate the performance of the proposed methods compared to existing methods. The proposed method based on the dependent Dirichlet process mixture can more accurately borrow from homogeneous historical controls while reducing the impact of heterogeneous historical controls compared to the typical Dirichlet process mixture. The proposed methods outperform existing methods in scenarios with heterogeneous historical controls, in which the meta-analytic approach is ineffective.

stat.ME

Sample size calculations for single-arm survival studies using transformations of the Kaplan-Meier estimator

In single-arm clinical trials with survival outcomes, the Kaplan-Meier estimator and its confidence interval are widely used to assess survival probability and median survival time. Since the asymptotic normality of the Kaplan-Meier estimator is a common result, the sample size calculation methods have not been studied in depth. An existing sample size calculation method is founded on the asymptotic normality of the Kaplan-Meier estimator using the log transformation. However, the small sample properties of the log transformed estimator are quite poor in small sample sizes (which are typical situations in single-arm trials), and the existing method uses an inappropriate standard normal approximation to calculate sample sizes. These issues can seriously influence the accuracy of results. In this paper, we propose alternative methods to determine sample sizes based on a valid standard normal approximation with several transformations that may give an accurate normal approximation even with small sample sizes. In numerical evaluations via simulations, some of the proposed methods provided more accurate results, and the empirical power of the proposed method with the arcsine square-root transformation tended to be closer to a prescribed power than the other transformations. These results were supported when methods were applied to data from three clinical trials.

stat.ME

Outlier detection and influence diagnostics in network meta-analysis

Network meta-analysis has been gaining prominence as an evidence synthesis method that enables the comprehensive synthesis and simultaneous comparison of multiple treatments. In many network meta-analyses, some of the constituent studies may have markedly different characteristics from the others, and may be influential enough to change the overall results. The inclusion of these "outlying" studies might lead to biases, yielding misleading results. In this article, we propose effective methods for detecting outlying and influential studies in a frequentist framework. In particular, we propose suitable influence measures for network meta-analysis models that involve missing outcomes and adjust the degree of freedoms appropriately. We propose three influential measures by a leave-one-trial-out cross-validation scheme: (1) comparison-specific studentized residual, (2) relative change measure for covariance matrix of the comparative effectiveness parameters, (3) relative change measure for heterogeneity covariance matrix. We also propose (4) a model-based approach using a likelihood ratio statistic by a mean-shifted outlier detection model. We illustrate the effectiveness of the proposed methods via applications to a network meta-analysis of antihypertensive drugs. Using the four proposed methods, we could detect three potential influential trials involving an obvious outlier that was retracted because of data falsifications. We also demonstrate that the overall results of comparative efficacy estimates and the ranking of drugs were altered by omitting these three influential studies.

stat.ME