Searcharxiv⌕ Search

arXiv subjects

Shunichiro Orihara

Publications and source records attributed to Shunichiro Orihara.

15 recordsLinked to original sources

Estimating the average treatment effect under limited overlap via Polynomial Approximation and Extrapolation

Estimating the average treatment effect (ATE) remains a fundamental challenge in observational studies in the presence of poor or limited covariate overlap. Although the inverse probability weighting (IPW) estimator is a widely used approach for estimating the ATE, its performance can deteriorate substantially when overlap is limited, often resulting in increased finite sample bias and unreliable confidence intervals. One common strategy is to shift attention from the original target estimand, the ATE, to alternative estimands such as a class of weighted ATEs that are less sensitive to extreme propensity scores; however, doing so changes the scientific question of interest. In this manuscript, we propose a novel ATE estimator that preserves the original target estimand, the ATE, while improving robustness to limited overlap. A key idea is that this class of estimands can be represented by a polynomial function of a hyperparameter characterizing the estimands. Exploiting this structure, the proposed method computes IPW estimators for a sequence of such estimands, models these estimates using a polynomial regression, and extrapolates to recover the ATE. We show that the estimator has consistency and asymptotic normality under weaker overlap conditions than required for the standard IPW estimator. Simulation studies demonstrate that the proposed method improves estimation accuracy and interval performance in settings with limited overlap. In addition to its theoretical and empirical advantages, the proposed approach has a clear interpretation and is easy to implement using standard statistical software.

stat.ME↗

Efficient Bayesian Inference in the Cox Model via Rank-Ordered Likelihood

In Bayesian inference for the Cox proportional hazards model, modeling the baseline hazard function is challenging. Recently, direct Bayesian inference using the partial likelihood is considered in the framework of general Bayesian inference. In terms of posterior computation, several studies have examined sampling algorithms under the Cox model. In this study, we propose two Gibbs sampling algorithms for Bayesian inference in the Cox proportional hazards model, motivated by a rank-ordered data representation and based on the Plackett--Luce and generalized Plackett--Luce models with P'{o}lya--Gamma data augmentation, referred to as PL-Cox and GPL-Cox, respectively. The two proposed methods offer practical advantages, as they do not require correction of posterior samples, naturally handle tied event times, and are readily extensible to shared frailty models. In simulation study, we considered multiple survival model settings, including continuous and discrete survival time models, as well as scenarios with varying degrees of ties, and found that the PL-Cox model exhibited relatively stable performance. In analyses of a large real dataset, the proposed methods remained computationally feasible, and the GPL-Cox model showed more favorable computational scalability than the PL-Cox model. In analyses of real data incorporating shared frailty, both methods demonstrated good computational efficiency.

stat.ME↗

On the Conservativeness of Robust Variance Estimators in Propensity Score Weighted Cox Models

In propensity score weighted analysis, robust variance that does not account for weight estimation is commonly used. In propensity score weighted Cox models (CoxPSW), the robust variance is known to be conservative when weights for the average treatment effect (ATE) are used, but it remains unclear whether this conservativeness also holds for other weighting schemes. This study evaluated the performance of the robust variance in CoxPSW when weights other than ATE are applied. We conducted an asymptotic comparison between the robust variance and a variance estimator that accounts for weight estimation under non-ATE weights. Their performance was further evaluated through simulation studies and real data analysis. The analytical results, simulations, and real data analysis indicated that the robust variance is not necessarily conservative in CoxPSW when weights other than ATE are used. These findings suggest that variance estimators that account for weight estimation should be used when applying non-ATE weights in CoxPSW.

stat.ME↗

Bayesian Doubly Robust Causal Inference via Posterior Coupling

Bayesian doubly robust (DR) causal inference faces a fundamental dilemma: joint modeling of outcome and propensity score suffers from the feedback problem where outcome information contaminates propensity score estimation, while two-step inference sacrifices valid posterior distributions for computational convenience. We resolve this dilemma through posterior coupling via entropic tilting. Our framework constructs independent posteriors for propensity score and outcome models, then couples them using entropic tilting to enforce the DR moment condition. This yields the first fully Bayesian DR estimator with an explicit posterior distribution. Theoretically, we establish three key properties: (i) when the outcome model is correctly specified, the tilted posterior coincides with the original; (ii) under propensity score model correctness, the posterior mean remains consistent despite outcome model misspecification; (iii) convergence rates improve for nonparametric outcome models. Simulations demonstrate superior bias reduction and efficiency compared to existing methods. We illustrate practical advantages of the proposed method through two applications: sensitivity analysis for unmeasured confounding in antihypertensive treatment effects on dementia, and high-dimensional confounder selection combining shrinkage priors with modified moment conditions for right heart catheterization mortality. We provide an R package implementing the proposed method.

stat.ME↗

Sample size re-estimation in blinded hybrid-control design using inverse probability weighting

With the increasing availability of data from historical studies and real-world data sources, hybrid control designs that incorporate external data into the evaluation of current studies are being increasingly adopted. In these designs, it is necessary to pre-specify during the planning phase the extent to which information will be borrowed from historical control data. However, if substantial differences in baseline covariate distributions between the current and historical studies are identified at the final analysis, the amount of effective borrowing may be limited, potentially resulting in lower actual power than originally targeted. In this paper, we propose two sample size re-estimation strategies that can be applied during the course of the blinded current study. Both strategies utilize inverse probability weighting (IPW) based on the probability of assignment to either the current or historical study. When large discrepancies in baseline covariates are detected, the proposed strategies adjust the sample size upward to prevent a loss of statistical power. The performance of the proposed strategies is evaluated through simulation studies, and their practical implementation is demonstrated using a case study based on two actual randomized clinical studies.

stat.ME↗

General Bayesian inference for causal effects using covariate balancing procedure

In observational studies, the propensity score plays a central role in estimating causal effects of interest. The inverse probability weighting (IPW) estimator is commonly used for this purpose. However, if the propensity score model is misspecified, the IPW estimator may produce biased estimates of causal effects. Previous studies have proposed some robust propensity score estimation procedures. However, these methods require considering parameters that dominate the uncertainty of sampling and treatment allocation. This study proposes a novel Bayesian estimating procedure that necessitates probabilistically deciding the parameter, rather than deterministically. Since the IPW estimator and propensity score estimator can be derived as solutions to certain loss functions, the general Bayesian paradigm, which does not require the considering the full likelihood, can be applied. Therefore, our proposed method only requires the same level of assumptions as ordinary causal inference contexts. The proposed Bayesian method demonstrates equal or superior results compared to some previous methods in simulation experimentss, and is also applied to real data, namely the Whitehall dataset.

stat.ME↗

Robust Estimation and Model Selection for the Controlled Directed Effect with Unmeasured Mediator-Outcome Confounders

Controlled Direct Effect (CDE) is one of the causal estimands used to evaluate both exposure and mediation effects on an outcome. When there are unmeasured confounders existing between the mediator and the outcome, the ordinary identification assumption does not work. In this manuscript, we consider an identification condition to identify CDE in the presence of unmeasured confounders. The key assumptions are: 1) the random allocation of the exposure, and 2) the existence of instrumental variables directly related to the mediator. Under these conditions, we propose a novel doubly robust estimation method, which work well if either the propensity score model or the baseline outcome model is correctly specified. Additionally, we propose a Generalized Information Criterion (GIC)-based model selection criterion for CDE that ensures model selection consistency. Our proposed procedure and related methods are applied to both simulation and real datasets to confirm the performance of these methods. Our proposed method can select the correct model with high probability and accurately estimate CDE.

stat.ME↗

Bayesian-based Propensity Score Subclassification Estimator

Subclassification estimators are one of the methods used to estimate causal effects of interest using the propensity score. This method is more stable compared to other weighting methods, such as inverse probability weighting estimators, in terms of the variance of the estimators. In subclassification estimators, the number of strata is traditionally set at five, and this number is not typically chosen based on data information. Even when the number of strata is selected, the uncertainty from the selection process is often not properly accounted for. In this study, we propose a novel Bayesian-based subclassification estimator that can assess the uncertainty in the number of strata, rather than selecting a single optimal number, using a Bayesian paradigm. To achieve this, we apply a general Bayesian procedure that does not rely on a likelihood function. This procedure allows us to avoid making strong assumptions about the outcome model, maintaining the same flexibility as traditional causal inference methods. With the proposed Bayesian procedure, it is expected that uncertainties from the design phase can be appropriately reflected in the analysis phase, which is sometimes overlooked in non-Bayesian contexts.

stat.ME↗

Nonparametric Bayesian Adjustment of Unmeasured Confounders in Cox Proportional Hazards Models

In observational studies, unmeasured confounders present a crucial challenge in accurately estimating desired causal effects. To calculate the hazard ratio (HR) in Cox proportional hazard models for time-to-event outcomes, two-stage residual inclusion and limited information maximum likelihood are typically employed. However, these methods are known to entail difficulty in terms of potential bias of HR estimates and parameter identification. This study introduces a novel nonparametric Bayesian method designed to estimate an unbiased HR, addressing concerns that previous research methods have had. Our proposed method consists of two phases: 1) detecting clusters based on the likelihood of the exposure and outcome variables, and 2) estimating the hazard ratio within each cluster. Although it is implicitly assumed that unmeasured confounders affect outcomes through cluster effects, our algorithm is well-suited for such data structures. The proposed Bayesian estimator has good performance compared with some competitors.

stat.ME↗

Evaluating the Conservativeness of Robust Sandwich Variance Estimator in Weighted Average Treatment Effects

In causal inference, the Inverse Probability Weighting (IPW) estimator is commonly used to estimate causal effects for estimands within the class of Weighted Average Treatment Effect (WATE). When constructing confidence intervals (CIs), robust sandwich variance estimators are frequently used for practical reasons. Although these estimators are easy to calculate using widely-used statistical software, they often yield narrow CIs for commonly applied estimands, such as the Average Treatment Effect on the Treated and the Average Treatment Effect for the Overlap Populations. In this manuscript, we reexamine the asymptotic variance of the IPW estimator and clarify the conditions under which CIs derived from the sandwich variance estimator are conservative. Additionally, we propose new criteria to assess the conservativeness of CIs. The results of this investigation are validated through simulation experiments and real data analysis.

stat.ME↗

Robust Estimating Method for Propensity Score Models and its Application to Some Causal Estimands: A review and proposal

In observational study, the propensity score has the central role to estimate causal effects. Since the propensity score is usually unknown, estimating by appropriate procedures is an indispensable step. A point to note that a causal effect estimator might have some bias if a propensity score model was misspecified; valid model construction is important. To overcome the problem, a variety of interesting methods has been proposed. In this paper, we review four methods: using ordinary logistic regression approach; CBPS proposed by Imai and Ratkovic; boosted CART proposed by McCaffrey and colleagues; a semiparametric strategy proposed by Liu and colleagues. Also, we propose the novel robust two step strategy: estimating each candidate model in the first step and integrating them in the second step. We confirm the performance of these methods through simulation examples by estimating the ATE and ATO proposed by Li and colleagues. From the results of the simulation examples, the boosted CART and CBPS with higher-order balancing condition have good properties; both the estimate of the ATE and ATO has the small variance and the absolute value of bias. The boosted CART and CBPS are useful for a variety of estimands and estimating procedures.

stat.ME↗

Frailty Model with Change Point for Survival Analysis

We propose a novel frailty model with change points applying random effects to a Cox proportional hazard model to adjust the heterogeneity between clusters. Because the frailty model includes random effects, the parameters are estimated using the expectation-maximization (EM) algorithm. Additionally, our model needs to estimate change points; we thus propose a new algorithm extending the conventional estimation algorithm to the frailty model with change points to solve the problem. We show a practical example to demonstrate how to estimate the change point and random effect. Our proposed model can be easily analyzed using the existing R package. We conducted simulation studies with three scenarios to confirm the performance of our proposed model. We re-analyzed data of two clinical trials to show the difference in analysis results with and without random effect. In conclusion, we confirmed that the frailty model with change points has a higher accuracy than the model without the random effect. Our proposed model is useful when heterogeneity needs to be taken into account. Additionally, the absence of heterogeneity did not affect the estimation of the regression coefficient parameters.

stat.ME↗

Valid Instrumental Variables Selection Methods using Negative Control Outcomes and Constructing Efficient Estimator

In observational studies, instrumental variable (IV) methods are commonly applied when there exists some unmeasured covariates. In Mendelian Randomization (MR), constructing an allele score by using many single nucleotide polymorphisms (SNPs) is often implemented; however, there are risks estimating biased causal effects by including some invalid IVs. Invalid IVs are candidates of IVs associated with some unobserved variables. To solve this problem, we propose a novel strategy in this paper: using Negative Control Outcomes (NCOs) as auxiliary variables. By using NCOs, we can essentialy select only valid IVs and exclude invalid IVs without any information of IV candidates. We also propose the new two-step estimating procedure and prove the semiparametric efficiency. We demonstrate the superior performance of the proposed estimator compared with existing estimators via simulation studies.

stat.ME↗

Likelihood-based Instrumental Variable Methods for Cox Proportional Hazard Models

In biometrics and related fields, the Cox proportional hazards model are widely used to analyze with covariate adjustment. However, when some covariates are not observed, an unbiased estimator usually cannot be obtained. Even if there are some unmeasured covariates, instrumental variable methods can be applied under some assumptions. In this paper, we propose the new instrumental variable estimator for the Cox proportional hazards model. The estimator is the similar feature as Martinez-Camblor et al., 2019, but not the same exactly; we use an idea of limited-information maximum likelihood. We show that the estimator has good theoretical properties. Also, we confirm properties of our method and previous methods through simulations datasets.

stat.ME↗

Limited-Information Maximum Likelihood based Model Selection Procedures for Binary Outcomes

Unmeasured covariates constitute one of the important problems in causal inference. Even if there are some unmeasured covariates, some instrumental variable methods such as a two-stage residual inclusion (2SRI) estimator, or a limited-information maximum likelihood (LIML) estimator can obtain an unbiased estimate for causal effects despite there being nonlinear outcomes such as binary outcomes; however, it requires that we specify not only a correct outcome model but also a correct treatment model. Therefore, detecting correct models is an important process. In this paper, we propose two model selection procedures: AIC-type and BIC-type, and confirm their properties. The proposed model selection procedures are based on a LIML estimator. We prove that a proposed BIC-type model selection procedure has model selection consistency, and confirm their properties of the proposed model selection procedures through simulation datasets.

stat.ME↗