SearcharxivSearch

arXiv subjects

Sho Komukai

Publications and source records attributed to Sho Komukai.

4 recordsLinked to original sources

Estimating the Average Treatment Effect under Limited Overlap via Polynomial Approximation and Extrapolation

Estimating the average treatment effect (ATE) remains a fundamental challenge in observational studies in the presence of poor or limited covariate overlap. Although the inverse probability weighting (IPW) estimator is a widely used approach for estimating the ATE, its performance can deteriorate substantially when overlap is limited, often resulting in increased finite sample bias and unreliable confidence intervals. One common strategy is to shift attention from the original target estimand, the ATE, to alternative estimands that are less sensitive to extreme propensity scores; however, doing so changes the scientific question of interest. In this manuscript, we propose a novel ATE estimator that preserves the original target estimand, the ATE, while improving robustness to limited overlap. A key idea is that a class of estimands can be expressed by a polynomial function of a hyperparameter characterizing the estimands. Exploiting this structure, the proposed method computes IPW estimators for a sequence of such estimands, models these estimates using a polynomial function, and extrapolates to recover the ATE. We show that the estimator has consistency and asymptotic normality under weaker overlap conditions than required for the standard IPW estimator. Simulation studies demonstrate that the proposed method improves estimation accuracy and interval performance in settings with limited overlap. In addition to its theoretical and empirical advantages, the proposed approach has a clear interpretation and is easy to implement using standard statistical software.

stat.ME

Data-Adaptive Integration with External Summary Data for Outcome Mean Estimation

Combining an internal individual-level study with readily available external summary statistics promises major efficiency gains at minimal additional cost, yet heterogeneity between sources can bias estimates for the internal target population. We develop a generalized entropy-balancing integration strategy that calibrates the internal individual-level sample to externally reported moments while retaining the internal population as the target, explicitly permitting a biased external sample. The weighted-regression version of our estimator is doubly robust: it remains consistent when either the outcome-regression model or the entropy-balancing model is correctly specified. When multiple balancing specifications are plausible, we introduce a data-adaptive entropy-family selection rule. For the final borrowing decision, we propose a bootstrap-based criterion comparing stabilized mean squared error (MSE) estimates for the selected entropy-balancing estimator and the internal sample mean. This criterion is selection consistent under fixed alternatives and reverts to the internal estimator when a nonvanishing bias is detected. Separately, under a linear homoscedastic benchmark, the asymptotic efficiency criteria admit geometric interpretations through the Mahalanobis distance and Pearson chi-squared divergence. The entropy-balancing estimators and numerical-experiment routines are implemented in the R package daisy. Simulations show stable MSE reductions for the weighted-regression estimator across calibrated distributional shifts and the predicted reversion toward the internal estimator under fixed simultaneous misspecification as the sample size increases. An application to nationwide public-access defibrillation records in Japan illustrates the resulting MSE-based borrowing decision.

stat.ME

On a fundamental problem in the analysis of cancer registry data

In epidemiology research with cancer registry data, it is often of primary interest to make inference on cancer death, not overall survival. Since cause of death is not easy to collect or is not necessarily reliable in cancer registries, some special methodologies have been introduced and widely used by using the concepts of the relative survival ratio and the net survival. In making inference of those measures, external life tables of the general population are utilized to adjust the impact of non-cancer death on overall survival. The validity of this adjustment relies on the assumption that mortality in the external life table approximates non-cancer mortality of cancer patients. However, the population used to calculate a life table may include cancer death and cancer patients. Sensitivity analysis proposed by Talbäck and Dickman to address it requires additional information which is often not easily available. We propose a method to make inference on the net survival accounting for potential presence of cancer patients and cancer death in the life table for the general population. The idea of adjustment is to consider correspondence of cancer mortality in the life table and that in the cancer registry. We realize a novel method to adjust cancer mortality in the cancer registry without any additional information to the standard analyses of cancer registries. Our simulation study revealed that the proposed method successfully removed the bias. We illustrate the proposed method with the cancer registry data in England.

stat.ME

Using clinical trial registries to inform Copas selection model for publication bias in meta-analysis

Prospective registration of study protocols in clinical trial registries is a useful way to minimize the risk of publication bias in meta-analysis, and several clinical trial registries are available nowadays. However, they are mainly used as a tool for searching studies and information submitted to the registries has not been utilized as efficiently as it could. In addressing publication bias in meta-analyses, sensitivity analysis with the Copas selection model is a more objective alternative to widely-used graphical methods such as the funnel-plot and the trim-and-fill method. Despite its ability to quantify the potential impact of publication bias, a drawback of the model is that some parameters not to be specified. This may result in some difficulty in interpreting the results of the sensitivity analysis. In this paper, we propose an alternative inference procedure for the Copas selection model by utilizing information from clinical trial registries. Our method provides a simple and accurate way to estimate all unknown parameters in the Copas selection model. A simulation study revealed that our proposed method resulted in smaller biases and more accurate confidence intervals than existing methods. Furthermore, two published meta-analyses had been re-analysed to demonstrate how to implement the proposed method in practice.

stat.ME