SearcharxivSearch

arXiv subjects

Hisashi Noma

Publications and source records attributed to Hisashi Noma.

At least 19 recordsLinked to original sources

Frequentist prediction intervals for random-effects meta-analysis via confidence-distribution propagation

Prediction intervals are increasingly recommended in random-effects meta-analysis because they describe the range of true effects expected in a future study or setting. Conventional frequentist intervals can have inadequate finite-sample coverage because uncertainty in the between-study variance is not fully propagated. We propose a confidence-distribution propagation method that carries uncertainty through the random-effects hierarchy. The method samples the between-study variance from a confidence distribution obtained by inverting the exact distribution of Cochran's Q and, conditional on each draw, samples the average effect from its corresponding normal confidence distribution before generating a future true effect. Prediction limits are empirical quantiles of the resulting Monte Carlo distribution. Across the scenarios examined, the proposed method improved or maintained coverage relative to existing frequentist intervals, including the Nagashima-Noma-Furukawa confidence-distribution bootstrap, with generally modest increases in expected width. The same Monte Carlo sample also yields confidence intervals for the average effect and heterogeneity measures. The method is implemented in the R package cdmeta available at CRAN (https://doi.org/10.32614/CRAN.package.cdmeta).

stat.ME

Target Trial Emulation with the R Package TTE: A Tutorial and Methodological Guide

Target trial emulation structures observational causal analyses around the protocol of an ideal randomized trial. By aligning eligibility, treatment assignment, time zero, and follow-up, it can reduce avoidable biases, but implementation still requires coordinated decisions about data construction, inverse probability weighting, diagnostics, outcome models, standardization, competing risks, and uncertainty estimation. This article provides a self-contained methodological guide and practical tutorial for TTE, an R package for target trial emulations with longitudinal observational data. We describe target trial protocols, intention-to-treat and per-protocol estimands, identification assumptions, baseline and person-period data structures, temporal ordering for longitudinal weights, stabilized treatment and censoring weights, weight truncation, balance and effective-sample-size diagnostics, weighted pooled discrete-time survival models, model-based standardization, competing-risk analysis, weighted Kaplan-Meier and Aalen-Johansen estimation, and cluster bootstrap at the original-individual level. Two fully synthetic examples illustrate end-to-end workflows: sodium-glucose cotransporter 2 inhibitor versus dipeptidyl peptidase-4 inhibitor initiation with all-cause death, and sequentially nested angiotensin receptor blocker versus calcium channel blocker trials with heart-failure hospitalization and competing death. The examples show how to obtain and diagnose estimates in R and interpret relative hazards, absolute risks, cumulative incidence, and differences between intention-to-treat and per-protocol effects.

stat.ME

Penalized likelihood inference for beta-binomial meta-analysis of proportions of rare events

Meta-analyses of proportions often involve sparse event counts and zero-event studies. The beta-binomial model has been used as a flexible random-effects model for pooling overdispersed and rare-event proportions. However, the ordinary maximum likelihood estimator (MLE) may suffer from finite-sample bias when few studies are available or the mean event probability is close to the boundary. In this article, we propose a maximum penalized likelihood estimator based on the Jeffreys-prior penalty, which has shown favorable finite-sample bias and stability properties in sparse-data models. We also develop Wald-type and profile penalized likelihood confidence intervals (CIs). In a simulation study, the proposed estimator generally achieved higher convergence rates, lower bias, and lower root mean squared error than the ordinary MLE, particularly with fewer studies or lower event probabilities. Additionally, the profile penalized likelihood CIs maintained coverage close to the nominal level. The proposed method provides a useful alternative for meta-analyses of rare-event proportions.

stat.ME

Robust inference methods of diagnostic test accuracy meta-analysis for influential outlying studies via density power divergence

In diagnostic test accuracy meta-analysis (DTA-MA), standard inference methods using bivariate random-effects models for jointly synthesizing sensitivity and specificity can be sensitive to outlying studies and may yield misleading conclusions. In this article, we propose frequentist outlier-robust statistical inference methods for DTA-MA based on density power divergence. The proposed methods automatically downweight influential outlying studies by modifying the estimating function using the robust divergence with a tuning parameter. To achieve robust yet statistically efficient inference in the presence of outlying studies, the proposed methods incorporate practical strategies for selecting the tuning parameter, including a data-adaptive criterion based on the Hyvärinen score. We also quantify the contributions of individual studies to the robust pooled estimates, facilitating interpretation of how outlying studies affect the results. We illustrate the effectiveness of the proposed methods through an application to a DTA-MA of the Mini-Mental State Examination. Simulation studies showed that the proposed methods reduced bias and root mean squared error relative to existing methods and improved coverage probability in the presence of outliers. The proposed methods enable a sensitivity analysis to assess whether the main results obtained using standard methods are driven by outlying studies.

stat.ME

Prediction intervals for random-effects meta-analysis: a confidence distribution approach

Prediction intervals are commonly used in meta-analysis with random-effects models. One widely used method, the Higgins-Thompson-Spiegelhalter prediction interval, replaces the heterogeneity parameter with its point estimate, but its validity strongly depends on a large sample approximation. This is a weakness in meta-analyses with few studies. We propose an alternative based on bootstrap and show by simulations that its coverage is close to the nominal level, unlike the Higgins-Thompson-Spiegelhalter method and its extensions. The proposed method was applied in three meta-analyses.

stat.ME

Influence analyses of "designs" for evaluating inconsistency in network meta-analysis

Network meta-analysis is an evidence synthesis method for comparing the effectiveness of multiple available treatments. To justify evidence synthesis, consistency is an important assumption; however, existing methods founded on statistical testing can be substantially limited in statistical power or have several drawbacks when handling multi-arm studies. Moreover, inconsistency can be theoretically explained as design-by-treatment interactions, and the primary purpose of such analyses is to prioritize the further investigation of specific "designs" to explore sources of bias and other issues that might influence the overall results. In this article, we propose an alternative framework for evaluating inconsistency using influence diagnostics methods, which enable the influence of individual designs on the overall results to be quantitatively evaluated. We provide four new methods, the averaged studentized residual, MDFFITS, Φ_d, and Ξ_d, to quantify the influence of individual designs through a "leave-one-design-out" analysis framework. We also propose a simple summary measure, the O-value, for prioritizing designs and interpreting these influential analyses in a straightforward manner. Furthermore, we propose another testing approach based on the leave-one-design-out analysis framework. By applying the new methods to a network meta-analysis of antihypertensive drugs and performing simulation studies, we demonstrate that the new methods accurately located potential sources of inconsistency. The proposed methods provide new insights into alternatives to existing test-based methods, especially the quantification of the influence of individual designs on the overall network meta-analysis results.

stat.ME

Ridge, lasso, and elastic-net estimations of the modified Poisson and least-squares regressions for binary outcome data

Logistic regression is a standard method in multivariate analysis for binary outcome data in epidemiological and clinical studies; however, the resultant odds-ratio estimates fail to provide directly interpretable effect measures. The modified Poisson and least-squares regressions are alternative standard methods that can provide risk-ratio and risk difference estimates without computational problems. However, the bias and invalid inference problems of these regression analyses under small or sparse data conditions (i.e.,the "separation" problem) have been insufficiently investigated. We show that the separation problem can adversely affect the inferences of the modified Poisson and least squares regressions, and to address these issues, we apply the ridge, lasso, and elastic-net estimating approaches to the two regression methods. As the methods are not founded on the maximum likelihood principle, we propose regularized quasi-likelihood approaches based on the estimating equations for these generalized linear models. The methods provide stable shrinkage estimates of risk ratios and risk differences even under separation conditions, and the lasso and elastic-net approaches enable simultaneous variable selection. We provide a bootstrap method to calculate the confidence intervals on the basis of the regularized quasi-likelihood estimation. The proposed methods are applied to a hematopoietic stem cell transplantation cohort study and the National Child Development Survey. We also provide an R package, regconfint, to implement these methods with simple commands.

stat.ME

MVPBT: R package for publication bias tests in meta-analysis of diagnostic accuracy studies

Meta-analysis for diagnostic test accuracy (DTA) has been a standard research method for synthesizing evidence from diagnostic studies. In DTA meta-analysis, although publication bias is an important source of bias, no certain methods similar to the Egger test in univariate meta-analysis have been developed to detect such bias. However, several recent studies have discussed these methods in the framework of multivariate meta-analysis, and some generalized Egger tests have been developed. The R package MVPBT (https://cran.r-project.org/web/packages/MVPBT/) was developed to implement the generalized Egger tests developed by Noma (2020; Biometrics 76, 1255-1259) for DTA meta-analysis. Noma's publication bias tests effectively incorporate the correlation information between multiple outcomes and are expected to improve the statistical powers. The present paper provides a nontechnical introduction and practical examples of data analyses of the publication bias tests of DTA meta-analysis using the MVPBT package.

stat.CO

Variance estimation for logistic regression in case-cohort studies

The logistic regression analysis proposed by Schouten et al. (Stat Med. 1993;12:1733-1745) has been a standard method in current statistical analysis of case-cohort studies, and it enables effective estimation of risk ratio from selected subsamples. Schouten et al. (1993) also proposed the standard error estimate of the risk ratio estimator can be calculated by the robust variance estimator. In this article, however, we show that the robust variance estimator does not account for the duplications of case and subcohort samples and generally has certain bias, i.e., inaccurate confidence intervals and P-values are possibly obtained. To address the invalid statistical inference problem, we provide an alternative bootstrap-based valid variance estimator. Through simulation studies, the bootstrap method consistently provided more precise confidence intervals compared with those provided by the robust variance method, while retaining adequate coverage probabilities. The conventional robust variance estimator has certain bias, and inadequate conclusions might be deduced. The bootstrap method would be an alternative effective approach in practice to provide accurate evidence.

stat.ME

pimeta: an R package of prediction intervals for random-effects meta-analysis

The prediction interval is gaining prominence in meta-analysis as it enables the assessment of uncertainties in treatment effects and heterogeneity between studies. However, coverage probabilities of the current standard method for constructing prediction intervals cannot retain their nominal levels in general, particularly when the number of synthesized studies is moderate or small, because their validities depend on large sample approximations. Recently, several methods have developed been to address this issue. This paper briefly summarizes the recent developments in methods of prediction intervals and provides readable examples using R for multiple types of data with simple code. The pimeta package is an R package that provides these improved methods to calculate accurate prediction intervals and graphical tools to illustrate these results. The pimeta package is listed in ``CRAN Task View: Meta-Analysis.'' The analysis is easily performed in R using a series of R packages.

stat.ME

Confidence intervals of prediction accuracy measures for multivariable prediction models based on the bootstrap-based optimism correction methods

In assessing prediction accuracy of multivariable prediction models, optimism corrections are essential for preventing biased results. However, in most published papers of clinical prediction models, the point estimates of the prediction accuracy measures are corrected by adequate bootstrap-based correction methods, but their confidence intervals are not corrected, e.g., the DeLong's confidence interval is usually used for assessing the C-statistic. These naive methods do not adjust for the optimism bias and do not account for statistical variability in the estimation of parameters in the prediction models. Therefore, their coverage probabilities of the true value of the prediction accuracy measure can be seriously below the nominal level (e.g., 95%). In this article, we provide two generic bootstrap methods, namely (1) location-shifted bootstrap confidence intervals and (2) two-stage bootstrap confidence intervals, that can be generally applied to the bootstrap-based optimism correction methods, i.e., the Harrell's bias correction, 0.632, and 0.632+ methods. In addition, they can be widely applied to various methods for prediction model development involving modern shrinkage methods such as the ridge and lasso regressions. Through numerical evaluations by simulations, the proposed confidence intervals showed favourable coverage performances. Besides, the current standard practices based on the optimism-uncorrected methods showed serious undercoverage properties. To avoid erroneous results, the optimism-uncorrected confidence intervals should not be used in practice, and the adjusted methods are recommended instead. We also developed the R package predboot for implementing these methods (https://github.com/nomahi/predboot). The effectiveness of the proposed methods are illustrated via applications to the GUSTO-I clinical trial.

stat.ME

Sample size calculations for single-arm survival studies using transformations of the Kaplan-Meier estimator

In single-arm clinical trials with survival outcomes, the Kaplan-Meier estimator and its confidence interval are widely used to assess survival probability and median survival time. Since the asymptotic normality of the Kaplan-Meier estimator is a common result, the sample size calculation methods have not been studied in depth. An existing sample size calculation method is founded on the asymptotic normality of the Kaplan-Meier estimator using the log transformation. However, the small sample properties of the log transformed estimator are quite poor in small sample sizes (which are typical situations in single-arm trials), and the existing method uses an inappropriate standard normal approximation to calculate sample sizes. These issues can seriously influence the accuracy of results. In this paper, we propose alternative methods to determine sample sizes based on a valid standard normal approximation with several transformations that may give an accurate normal approximation even with small sample sizes. In numerical evaluations via simulations, some of the proposed methods provided more accurate results, and the empirical power of the proposed method with the arcsine square-root transformation tended to be closer to a prescribed power than the other transformations. These results were supported when methods were applied to data from three clinical trials.

stat.ME

Flexible random-effects distribution models for meta-analysis

In meta-analysis, the random-effects models are standard tools to address between-study heterogeneity in evidence synthesis analyses. For the random-effects distribution models, the normal distribution model has been adopted in most systematic reviews due to its computational and conceptual simplicity. However, the restrictive model assumption might have serious influences on the overall conclusions in practices. In this article, we first provide two examples of real-world evidence that clearly show that the normal distribution assumption is unsuitable. To address the model restriction problem, we propose alternative flexible random-effects models that can flexibly regulate skewness, kurtosis and tailweight: skew normal distribution, skew t-distribution, asymmetric Subbotin distribution, Jones-Faddy distribution, and sinh-arcsinh distribution. We also developed a R package, flexmeta, that can easily perform these methods. Using the flexible random-effects distribution models, the results of the two meta-analyses were markedly altered, potentially influencing the overall conclusions of these systematic reviews. The flexible methods and computational tools can provide more precise evidence, and these methods would be recommended at least as sensitivity analysis tools to assess the influence of the normal distribution assumption of the random-effects model.

stat.AP

Confidence interval for the AUC of SROC curve and some related methods using bootstrap for meta-analysis of diagnostic accuracy studies

The area under the curve (AUC) of summary receiver operating characteristic (SROC) curve is a primary statistical outcome for meta-analysis of diagnostic test accuracy studies (DTA). However, its confidence interval has not been reported in most of DTA meta-analyses, because no certain methods and statistical packages have been provided. In this article, we provide a bootstrap algorithm for computing the confidence interval of the AUC. Also, using the bootstrap framework, we can conduct a bootstrap test for assessing significance of the difference of AUCs for multiple diagnostic tests. In addition, we provide an influence diagnostic method based on the AUC by leave-one-study-out analyses. We present illustrative examples using two DTA met-analyses for diagnostic tests of cervical cancer and asthma. We also developed an easy-to-handle R package dmetatools for these computations. The various quantitative evidence provided by these methods certainly supports the interpretations and precise evaluations of statistical evidence of DTA meta-analyses.

stat.AP

Efficient testing and effect size estimation for set-based genetic association inference via semiparametric multilevel mixture modeling: Application to a genome-wide association study of coronary artery disease

In genetic association studies, rare variants with extremely small allele frequency play a crucial role in complex traits, and the set-based testing methods that jointly assess the effects of groups of single nucleotide polymorphisms (SNPs) were developed to improve powers for the association tests. However, the powers of these tests are still severely limited due to the extremely small allele frequency, and precise estimations for the effect sizes of individual SNPs are substantially impossible. In this article, we provide an efficient set-based inference framework that addresses the two important issues simultaneously based on a Bayesian semiparametric multilevel mixture model. We propose to use the multilevel hierarchical model that incorporate the variations in set-specific effects and variant-specific effects, and to apply the optimal discovery procedure (ODP) that achieves the largest overall power in multiple significance testing. In addition, we provide Bayesian optimal "set-based" estimator of the empirical distribution of effect sizes. Efficiency of the proposed methods is demonstrated through application to a genome-wide association study of coronary artery disease (CAD), and through simulation studies. These results suggested there could be a lot of rare variants with large effect sizes for CAD, and the number of significant sets detected by the ODP was much greater than those by existing methods.

stat.ME

Re-evaluation of the comparative effectiveness of bootstrap-based optimism correction methods in the development of multivariable clinical prediction models

Multivariable predictive models are important statistical tools for providing synthetic diagnosis and prognostic algorithms based on multiple patients' characteristics. Their apparent discriminant and calibration measures usually have overestimation biases (known as 'optimism') relative to the actual performances for external populations. Existing statistical evidence and guidelines suggest that three bootstrap-based bias correction methods are preferable in practice, namely Harrell's bias correction and the .632 and .632+ estimators. Although Harrell's method has been widely adopted in clinical studies, simulation-based evidence indicates that the .632+ estimator may perform better than the other two methods. However, there is limited evidence and these methods' actual comparative effectiveness is still unclear. In this article, we conducted extensive simulations to compare the effectiveness of these methods, particularly using the following modern regression models: conventional logistic regression, stepwise variable selections, Firth's penalized likelihood method, ridge, lasso, and elastic-net. Under relatively large sample settings, the three bootstrap-based methods were comparable and performed well. However, all three methods had biases under small sample settings, and the directions and sizes of the biases were inconsistent. In general, the .632+ estimator is recommended, but we provide several notes concerning the operating characteristics of each method.

stat.AP

Efficient screening of predictive biomarkers for individual treatment selection

The development of molecular diagnostic tools to achieve individualized medicine requires identifying predictive biomarkers associated with subgroups of individuals who might receive beneficial or harmful effects from different available treatments. However, due to the large number of candidate biomarkers in the large-scale genetic and molecular studies, and complex relationships among clinical outcome, biomarkers and treatments, the ordinary statistical tests for the interactions between treatments and covariates have difficulties from their limited statistical powers. In this paper, we propose an efficient method for detecting predictive biomarkers. We employ weighted loss functions of Chen et al. (2017) to directly estimate individual treatment scores and propose synthetic posterior inference for effect sizes of biomarkers. We develop an empirical Bayes approach, namely, we estimate unknown hyperparameters in the prior distribution based on data. We then provide efficient screening methods for the candidate biomarkers via optimal discovery procedure with adequate control of false discovery rate. The proposed method is demonstrated in simulation studies and an application to a breast cancer clinical study in which the proposed method was shown to detect the much larger numbers of significant biomarkers than existing standard methods.

stat.ME

Outlier detection and influence diagnostics in network meta-analysis

Network meta-analysis has been gaining prominence as an evidence synthesis method that enables the comprehensive synthesis and simultaneous comparison of multiple treatments. In many network meta-analyses, some of the constituent studies may have markedly different characteristics from the others, and may be influential enough to change the overall results. The inclusion of these "outlying" studies might lead to biases, yielding misleading results. In this article, we propose effective methods for detecting outlying and influential studies in a frequentist framework. In particular, we propose suitable influence measures for network meta-analysis models that involve missing outcomes and adjust the degree of freedoms appropriately. We propose three influential measures by a leave-one-trial-out cross-validation scheme: (1) comparison-specific studentized residual, (2) relative change measure for covariance matrix of the comparative effectiveness parameters, (3) relative change measure for heterogeneity covariance matrix. We also propose (4) a model-based approach using a likelihood ratio statistic by a mean-shifted outlier detection model. We illustrate the effectiveness of the proposed methods via applications to a network meta-analysis of antihypertensive drugs. Using the four proposed methods, we could detect three potential influential trials involving an obvious outlier that was retracted because of data falsifications. We also demonstrate that the overall results of comparative efficacy estimates and the ranking of drugs were altered by omitting these three influential studies.

stat.ME