Searcharxiv⌕ Search

arXiv subjects

Tomotaka Momozaki

Publications and source records attributed to Tomotaka Momozaki.

13 recordsLinked to original sources

Semiparametric Copula Estimation for Spatially Correlated Multivariate Mixed Outcomes: Analyzing Visual Sightings of Fin Whales from a Line Transect Survey

For marine biologists, ascertaining the dependence structures between marine species and marine environments, such as sea surface temperature and ocean depth, is imperative for defining ecosystem functioning and providing insights into the dynamics of marine ecosystems. However, obtained data include not only continuous but also discrete data, such as binaries and counts (referred to as mixed outcomes), as well as spatial correlations, both of which make conventional multivariate analysis tools impractical. To solve this issue, we propose semiparametric Bayesian inference and develop an efficient algorithm for computing the posterior of the dependence structure based on the rank likelihood under a latent multivariate spatial Gaussian process using the Markov chain Monte Carlo method. To alleviate the computational intractability caused by the Gaussian process, we also provide a scalable implementation that leverages the nearest-neighbor Gaussian process. Extensive numerical experiments reveal that the proposed method reliably infers the dependence structures of spatially correlated mixed outcomes. Finally, we apply the proposed method to a dataset collected during an international synoptic krill survey in the Scotia Sea of the Antarctic Peninsula to infer the dependence structure between fin whales (Balaenoptera physalus), krill biomass, and relevant oceanographic data.

stat.ME↗

Convergence fragility in probit Bayesian kernel machine regression implemented in the bkmr R package for binary-outcome environmental mixture analyses: a simulation study

Background. Bayesian kernel machine regression (BKMR) is widely used for exposure-mixture analyses with binary outcomes through a probit extension. Because a bkmr fit can complete without providing adequate effective posterior information, simulation studies should separate execution success from MCMC convergence diagnostics. Methods. We evaluated the public bkmr probit workflow using bkmr::SimData() for data generation, bkmr::kmbayes() for model fitting, and posterior for convergence diagnostics. The balanced generator used family = "binomial", hfun = 2, beta.true = 0.5, ind = 1:2, and M = 4. SimData() generated the covariate as X = 3*cos(z1) + 2*rnorm(n). Four chains were initialized with chain-specific randomized starting values generated reproducibly from the fixed initial-value base seed 20260621. These values affected only the initial state of the sampler and did not alter the BKMR model, default priors, or Metropolis-Hastings proposals. Results. Of 431 prespecified tasks, 430 returned fitted objects and one task had a numerical non-completion. Diagnostic adequacy was limited: rank-normalized R-hat <= 1.01 threshold was achieved in 55/431 tasks, bulk-ESS >= 400 in 85/431, tail-ESS >= 400 in 44/431, and both ESS criteria in 44/431. The primary diagnostic criterion, R-hat at or below the 1.01 threshold with both bulk-ESS and tail-ESS >= 400, was met in 30/431 prespecified tasks, corresponding to 30/430 completed fits. Conclusions. Completion of probit BKMR fits in bkmr should not be equated with convergence of the retained MCMC draws. Applied analyses should report the number of chains, warmup and retained iterations, rank-normalized R-hat, bulk-ESS, and tail-ESS rather than rely on a fixed iteration count or on fit completion alone.

stat.AP↗

Bayesian Doubly Robust Causal Inference via Posterior Coupling

Bayesian doubly robust (DR) causal inference faces a fundamental dilemma: joint modeling of outcome and propensity score suffers from the feedback problem where outcome information contaminates propensity score estimation, while two-step inference sacrifices valid posterior distributions for computational convenience. We resolve this dilemma through posterior coupling via entropic tilting. Our framework constructs independent posteriors for propensity score and outcome models, then couples them using entropic tilting to enforce the DR moment condition. This yields the first fully Bayesian DR estimator with an explicit posterior distribution. Theoretically, we establish three key properties: (i) when the outcome model is correctly specified, the tilted posterior coincides with the original; (ii) under propensity score model correctness, the posterior mean remains consistent despite outcome model misspecification; (iii) convergence rates improve for nonparametric outcome models. Simulations demonstrate superior bias reduction and efficiency compared to existing methods. We illustrate practical advantages of the proposed method through two applications: sensitivity analysis for unmeasured confounding in antihypertensive treatment effects on dementia, and high-dimensional confounder selection combining shrinkage priors with modified moment conditions for right heart catheterization mortality. We provide an R package implementing the proposed method.

stat.ME↗

Sampling from density power divergence-based generalized posterior distribution via stochastic optimization

Robust Bayesian inference using density power divergence (DPD) has emerged as a promising approach for handling outliers in statistical estimation. Although the DPD-based posterior offers theoretical guarantees of robustness, its practical implementation faces significant computational challenges, particularly for general parametric models with intractable integral terms. These challenges are specifically pronounced in high-dimensional settings, where traditional numerical integration methods are inadequate and computationally expensive. Herein, we propose a novel {approximate} sampling methodology that addresses these limitations by integrating the loss-likelihood bootstrap with a stochastic gradient descent algorithm specifically designed for DPD-based estimation. Our approach enables efficient and scalable sampling from DPD-based posteriors for a broad class of parametric models, including those with intractable integrals. We further extend it to accommodate generalized linear models. Through comprehensive simulation studies, we demonstrate that our method efficiently samples from DPD-based posteriors, offering superior computational scalability compared to conventional methods, specifically in high-dimensional settings. The results also highlight its ability to handle complex parametric models with intractable integral terms.

stat.ME↗

Improving Maximum Tolerated Dose Selection in Model-Assisted Designs for Phase I Trials through Bayesian Dose-Response Model

Model-assisted designs have garnered significant attention in recent years due to their high accuracy in identifying the maximum tolerated dose (MTD) and their operational simplicity. To identify the MTD, they employ estimated dose limiting toxicity (DLT) probabilities via isotonic regression with pool-adjacent violators algorithm (PAVA) after trials have been completed. PAVA adjusts independently estimated DLT probabilities with the Bayesian binomial model at each dose level using posterior variances ensure the monotonicity that toxicity increases with dose. However, in small sample settings such as Phase I oncology trials, this approach can lead to unstable DLT probability estimates and reduce MTD selection accuracy. To address this problem, we propose a novel MTD identification strategy in model-assisted designs that leverages a Bayesian dose-response model. Employing the dose-response model allows for stable estimation of the DLT probabilities under the monotonicity by borrowing information across dose levels, leading to an improvement in MTD identification accuracy. We discuss the specification of prior distributions that can incorporate information from similar trials or the absence of such information. We examine dose-response models employing logit, log-log, and complementary log-log link functions to assess the impact of link function differences on the accuracy of MTD selection. Through extensive simulations, we demonstrate that the proposed approach improves MTD selection accuracy by more than 10\% in some scenarios and by approximately 6\% on average compared to conventional approach. These findings indicate that the proposed approach can contribute to further enhancing the efficiency of Phase I oncology trials.

stat.AP↗

Robust Spatio-Temporal Distributional Regression

Motivated by investigating spatio-temporal patterns of the distribution of continuous variables, we consider describing the conditional distribution function of the response variable incorporating spatio-temporal components given predictors. In many applications, continuous variables are observed only as threshold-categorized data due to measurement constraints. For instance, ecological measurements often categorize sizes into intervals rather than recording exact values due to practical limitations. To recover the conditional distribution function of the underlying continuous variables, we consider a distribution regression employing models for binomial data obtained at each threshold value. However, depending on spatio-temporal conditions and predictors, the distribution function may frequently exhibit boundary values (zero or one), which can occur either structurally or randomly. This makes standard binomial models inadequate, requiring more flexible modeling approaches. To address this issue, we propose a boundary-inflated binomial model incorporating spatio-temporal components. The model is a three-component mixture of the binomial model and two Dirac measures at zero and one. We develop a computationally efficient Bayesian inference algorithm using Pólya-Gamma data augmentation and dynamic Gaussian predictive processes. Extensive simulation experiments demonstrate that our procedure significantly outperforms distribution regression methods based on standard binomial models across various scenarios.

stat.ME↗

General Bayesian inference for causal effects using covariate balancing procedure

In observational studies, the propensity score plays a central role in estimating causal effects of interest. The inverse probability weighting (IPW) estimator is commonly used for this purpose. However, if the propensity score model is misspecified, the IPW estimator may produce biased estimates of causal effects. Previous studies have proposed some robust propensity score estimation procedures. However, these methods require considering parameters that dominate the uncertainty of sampling and treatment allocation. This study proposes a novel Bayesian estimating procedure that necessitates probabilistically deciding the parameter, rather than deterministically. Since the IPW estimator and propensity score estimator can be derived as solutions to certain loss functions, the general Bayesian paradigm, which does not require the considering the full likelihood, can be applied. Therefore, our proposed method only requires the same level of assumptions as ordinary causal inference contexts. The proposed Bayesian method demonstrates equal or superior results compared to some previous methods in simulation experimentss, and is also applied to real data, namely the Whitehall dataset.

stat.ME↗

On window mean survival time with interval-censored data

In recent years, cancer clinical trials have increasingly encountered non proportional hazards (NPH) scenarios, particularly with the emergence of immunotherapy. In randomized controlled trials comparing immunotherapy with conventional chemotherapy or placebo, late difference and early crossing survivals scenarios are commonly observed. In such cases, window mean survival time (WMST), the area under the survival curve within a pre-specified interval $[τ_0, τ_1]$, has gained increasing attention due to its superior power compared to restricted mean survival time (RMST), the area under the survival curve up to a pre-specified time point. Considering the increasing use of progression-free survival as a co-primary endpoint alongside overall survival, there is a critical need to establish a WMST estimation method for interval-censored data; however, sufficient research has yet to be conducted. To bridge this gap, this study proposes a WMST inference method utilizing one-point imputations and Turnbull's method. Extensive numerical simulations demonstrate that the WMST estimation method using mid-point imputation for interval-censored data exhibits comparable performance to that using Turnbull's method. Since the former facilitates standard error calculation, we adopt it as the standard method. Numerical simulations on two-sample tests confirm that the proposed WMST testing method have higher power than RMST in late difference and early crossing survival scenarios, while having compatible power to the log-rank test under the PH. Furthermore, even when pre-specified $τ_0$ deviated from the clinically desirable time point, WMST consistently maintains higher power than RMST in late difference and early crossing survivals scenarios.

stat.AP↗

Bayesian-based Propensity Score Subclassification Estimator

Subclassification estimators are one of the methods used to estimate causal effects of interest using the propensity score. This method is more stable compared to other weighting methods, such as inverse probability weighting estimators, in terms of the variance of the estimators. In subclassification estimators, the number of strata is traditionally set at five, and this number is not typically chosen based on data information. Even when the number of strata is selected, the uncertainty from the selection process is often not properly accounted for. In this study, we propose a novel Bayesian-based subclassification estimator that can assess the uncertainty in the number of strata, rather than selecting a single optimal number, using a Bayesian paradigm. To achieve this, we apply a general Bayesian procedure that does not rely on a likelihood function. This procedure allows us to avoid making strong assumptions about the outcome model, maintaining the same flexibility as traditional causal inference methods. With the proposed Bayesian procedure, it is expected that uncertainties from the design phase can be appropriately reflected in the analysis phase, which is sometimes overlooked in non-Bayesian contexts.

stat.ME↗

Improving the accuracy of estimating indexes in contingency tables using Bayesian estimators

In contingency table analysis, one is interested in testing whether a model of interest (e.g., the independent or symmetry model) holds using goodness-of-fit tests. When the null hypothesis where the model is true is rejected, the interest turns to the degree to which the probability structure of the contingency table deviates from the model. Many indexes have been studied to measure the degree of the departure, such as the Yule coefficient and Cramér coefficient for the independence model, and Tomizawa's symmetry index for the symmetry model. The inference of these indexes is performed using sample proportions, which are estimates of cell probabilities, but it is well-known that the bias and mean square error (MSE) values become large without a sufficient number of samples. To address the problem, this study proposes a new estimator for indexes using Bayesian estimators of cell probabilities. Assuming the Dirichlet distribution for the prior of cell probabilities, we asymptotically evaluate the value of MSE when plugging the posterior means of cell probabilities into the index, and propose an estimator of the index using the Dirichlet hyperparameter that minimizes the value. Numerical experiments show that when the number of samples per cell is small, the proposed method has smaller values of bias and MSE than other methods of correcting estimation accuracy. We also show that the values of bias and MSE are smaller than those obtained by using the uniform and Jeffreys priors.

stat.ME↗

Generalized Cramér's coefficient via $f$-divergence for contingency tables

Various measures in two-way contingency table analysis have been proposed to express the strength of association between row and column variables in contingency tables. Tomizawa et al. (2004) proposed more general measures, including Cramér's coefficient, using the power-divergence. In this paper, we propose measures using the $f$-divergence that has a wider class than the power-divergence. Unlike statistical hypothesis tests, these measures provide quantification of the association structure in contingency tables. The contribution of our study is proving that a measure applying a function that satisfies the condition of the $f$-divergence has desirable properties for measuring the strength of association in contingency tables. With this contribution, we can easily construct a new measure using a divergence that has essential properties for the analyst. For example, we conducted numerical experiments with a measure applying the $θ$-divergence. Furthermore, we can give further interpretation of the association between the row and column variables in the contingency table, which could not be obtained with the conventional one. We also show a relationship between our proposed measures and the correlation coefficient in the bivariate normal distribution of latent variables in the contingency tables.

stat.ME↗

Robustness of Bayesian ordinal response model against outliers via divergence approach

Ordinal response model is a popular and commonly used regression for ordered categorical data in a wide range of fields such as medicine and social sciences. However, it is empirically known that the existence of ``outliers'', combinations of the ordered categorical response and covariates that are heterogeneous compared to other pairs, makes the inference with the ordinal response model unreliable. In this article, we prove that the posterior distribution in the ordinal response model does not satisfy the posterior robustness with any link functions, i.e., the posterior cannot ignore the influence of large outliers. Furthermore, to achieve robust Bayesian inference in the ordinal response model, this article defines general posteriors in the ordinal response model with two robust divergences (the density-power and $γ$-divergences) based on the framework of the general posterior inference. We also provide an algorithm for generating posterior samples from the proposed posteriors. The robustness of the proposed methods against outliers is clarified from the posterior robustness and the index of robustness based on the Fisher-Rao metric. Through numerical experiments on artificial data and two real datasets, we show that the proposed methods perform better than the ordinary bayesian methods with and without outliers in the data for various link functions.

stat.ME↗

Robustness against outliers in ordinal response model via divergence approach

This study deals with the problem of outliers in ordinal response model, which is a regression on ordered categorical data as the response variable. ``Outlier" means that the combination of ordered categorical data and its covariates is heterogeneous compared to other pairs. Although the ordinal response model is important for data analysis in various fields such as medicine and social sciences, it is known that the maximum likelihood method with probit, logit, log-log and complementary log-log link functions, which are often used, is strongly affected by outliers, and statistical analysts are forced to limit their analysis when there may be outliers in the data. To solve this problem, this paper provides inference methods with two robust divergences (the density-power and $γ$-divergences). We also derive influence functions for the proposed methods and show conditions on the link function for them to be bounded and to redescendence. Since the commonly used link functions satisfy these conditions, the analyst can perform robust and flexible analysis with our methods. In addition, and this is a result that further highlights our contributions, we show that the influence function in the maximum likelihood method does not have redescendence for any link function in the ordinal response model. Through numerical experiments using artificial and two real data, we show that the proposed methods perform better than the maximum likelihood method with and without outliers in the data for various link functions.

stat.ME↗