SearcharxivSearch

arXiv subjects

Shonosuke Sugasawa

Publications and source records attributed to Shonosuke Sugasawa.

At least 19 recordsLinked to original sources

Beyond Tweedie's Formula: Conditional Score Modeling for Empirical Bayes Inference

We propose conditional f-modeling (Cf-modeling), a framework for empirical Bayes inference with covariates. A central identity shows that the conditional marginal score function determines not only the posterior mean through Tweedie's formula, but also the posterior moment-generating function, providing a basis for recovering posterior quantities without explicit prior modeling. Motivated by this observation, we treat the conditional marginal score as the primary object of inference and estimate it directly using an energy-based representation and Hyv\"arinen score matching, thereby avoiding potentially intractable covariate-dependent normalizing constants. The resulting framework flexibly accommodates covariate effects and heteroscedasticity and provides a practical approach to posterior moment estimation and uncertainty quantification. We demonstrate the effectiveness of the proposed method through simulations and an RNA-seq application.

stat.ME

Finite Mixtures of Generalized Estimating Equations for Clustering Multivariate Correlated Outcomes

Multivariate correlated outcomes occur across disciplines, including ecology, social sciences, and psychometrics. This paper focuses on clustering these outcomes across observational units, specifically, finding groups of units with the same ``outcome profile". Our motivation comes from bioregionalization in ecology, which aims to cluster sites into bioregions with the same species profiles, where site membership can depend on environmental or habitat covariates. To accomplish this, we propose finite mixtures of generalized estimating equations (MixGEE). Unlike existing approaches to model-based bioregionalization, MixGEE partitions sites into regions while accounting for between-species correlations through a region-specific working correlation structure. Thus, each region is characterized by a marginal species mean vector and a between-species correlation matrix. Unlike likelihood-based finite mixture models, MixGEE does not require a full joint distribution of the multivariate outcomes. Instead, we construct a pseudo-posterior probability for region membership motivated by the large-sample distribution of the estimating equation. This leads to an iterative algorithm alternating between updating these probabilities and solving weighted estimating equations. We determine the number of regions using cross-validation based on predictive performance for held-out sites and species components, and use a clustered Dirichlet random-weight bootstrap for uncertainty quantification. Simulations demonstrate reliable estimation and inference under various correlation structures and more stable selection of the number of groups than methods that ignore dependence. Applying MixGEE to presence--absence records of fish species around the Kerguelen Plateau reveals three distinct fish assemblage profiles with heterogeneous occurrence and within-site correlation patterns.

stat.ME

Log-regularly varying scale mixture of asymmetric Laplaces for robust Bayesian quantile regression

Bayesian quantile regression based on the asymmetric Laplace (AL) distribution can be sensitive to extreme observations because of its exponentially decaying tails. We propose a robust error distribution constructed as a finite mixture of the AL distribution and a log-Pareto scale mixture of asymmetric Laplace distributions (LPAL). Unlike a direct log-Pareto extension of the normal location-scale representation of the AL distribution, the proposed AL-LPAL mixture preserves the prescribed quantile and exhibits log-regularly varying behavior in both tails. The LPAL component also has an unbounded density at the target quantile, yielding a distribution that combines sharp central concentration with super-heavy tails. We establish posterior robustness under arbitrarily extreme contamination and provide sufficient conditions for the existence of posterior moments of the regression coefficients and scale parameter. For posterior computation, we develop a Gibbs sampler using latent-variable augmentations and a computationally efficient mean-field variational Bayes approximation. Simulation studies show that the proposed method is competitive under moderate contamination and maintains stable point estimation with comparatively concentrated posterior intervals, particularly when severe contamination affects the quantile of interest. Applications to carbon dioxide and Boston housing data, using the same preprocessing as existing robust Bayesian quantile regression analyses, show favorable predictive performance across nearly all quantile levels and loss criteria considered.

stat.ME

Small Area Estimation under Spatial Regimes: Spatially Clustered Fay-Herriot Models for Agricultural Indicators

Area-level small area estimation (SAE) models, such as the Fay--Herriot (FH) model, borrow strength across domains through covariates and random effects, but they can struggle when the relationship between the covariates and the outcome is spatially heterogeneous, that is, when it changes across the spatial domain of interest. We propose a spatially-clustered FH (SC-FH) framework that simultaneously (i) estimates cluster-specific regression coefficients and random effects variances and (ii) generates spatially coherent partitions of the geographical domain. Estimation maximizes a penalized likelihood that augments the FH likelihood with a Potts-type spatial cohesion term over the areal adjacency graph, through an efficient strategy that alternates between sequential label updates and closed-form FH updates within clusters. In simulation experiments run on the real geography of the application, the method recovers the latent regimes almost exactly whenever they are separated in the covariate--response space and improves prediction accuracy over the standard FH benchmark, with the spatial penalty acting as a stabilizer of both classification and estimation. An empirical application to the average standard output of farms in the Po Valley (Northern Italy) identifies two spatially compact production regimes with significantly different cluster-wise coefficients, and shows that the clusterwise predictor improves on the direct estimates while avoiding the over-shrinkage of the pooled model.

stat.AP

Information Gap and Feasibility-Aware Inference in Binomial Logistic Mixtures

This paper studies the information gap between mixture detection and label recovery in binomial logistic mixtures. Standard likelihood-based criteria such as the Bayesian information criterion (BIC) can detect the presence of two components, but this does not guarantee that the corresponding labels are recoverable. We show that this gap is intrinsic to binomial logistic mixtures with a fixed number of trials: observed-data evidence for mixture structure and per-observation information for label recovery have different local orders in the component separation, and only the former accumulates with the sample size. As a result, there exists a detectable-but-unrecoverable regime in which BIC selects two components while the posterior labels remain essentially uninformative. To address this issue, we propose two feasibility-aware inference procedures: a recoverability-aware BIC with a posterior-entropy penalty and an entropy-regularized estimator that mitigates the tendency of the maximum likelihood estimator to produce overly separated components and overly concentrated posterior responsibilities. Numerical experiments confirm the predicted gap and demonstrate that the proposed methods avoid misleading component selections and improve the calibration of posterior label probabilities.

stat.ML

Causal Small Area Estimation with Survey-only Covariates

Area-specific causal inference is important in many policy and survey applications, where the goal is to evaluate treatment effects for small geographic or demographic domains. Existing causal small area estimation methods, however, typically rely on a strong data requirement that treatment status is observed for all units in the population. This assumption is often unrealistic in practical survey settings, where both treatment and outcome variables are observed only for sampled units, while auxiliary covariates are available for the full population. To address this limitation, we develop a new identification strategy for area-specific treatment effects under this more realistic data structure by combining survey-only covariates with population-level auxiliary information. Based on this result, we propose a doubly robust estimator that remains consistent when either the outcome regression model or the treatment and area assignment models are correctly specified. We further derive the semiparametric efficiency bound for the target parameter and show that the proposed estimator attains this bound under regularity conditions. Simulation studies demonstrate favorable finite-sample performance, particularly in settings with small sample sizes within areas, and an empirical application illustrates the practical relevance of the proposed framework.

math.ST

Efficient Bayesian Inference in the Cox Model via Rank-Ordered Likelihood

In Bayesian inference for the Cox proportional hazards model, modeling the baseline hazard function is challenging. Recently, direct Bayesian inference using the partial likelihood is considered in the framework of general Bayesian inference. In terms of posterior computation, several studies have examined sampling algorithms under the Cox model. In this study, we propose two Gibbs sampling algorithms for Bayesian inference in the Cox proportional hazards model, motivated by a rank-ordered data representation and based on the Plackett--Luce and generalized Plackett--Luce models with P'{o}lya--Gamma data augmentation, referred to as PL-Cox and GPL-Cox, respectively. The two proposed methods offer practical advantages, as they do not require correction of posterior samples, naturally handle tied event times, and are readily extensible to shared frailty models. In simulation study, we considered multiple survival model settings, including continuous and discrete survival time models, as well as scenarios with varying degrees of ties, and found that the PL-Cox model exhibited relatively stable performance. In analyses of a large real dataset, the proposed methods remained computationally feasible, and the GPL-Cox model showed more favorable computational scalability than the PL-Cox model. In analyses of real data incorporating shared frailty, both methods demonstrated good computational efficiency.

stat.ME

Dynamic Bayesian regression quantile synthesis for forecasting outlook-at-risk

This paper proposes dynamic Bayesian regression quantile synthesis (DRQS), a novel method for quantile forecasting within the Bayesian predictive synthesis (BPS) framework designed to combine quantile-specific information from multiple agent models. While existing BPS approaches primarily focus on mean forecasting, our method directly targets the conditional quantiles of the response variable by utilizing the asymmetric Laplace distribution for the synthesis function. The resulting framework can be interpreted as a dynamic quantile linear model with latent predictors. We extend the univariate DRQS to a multivariate setting-factor DRQS (FDRQS)-by introducing a time-varying latent factor structure for the synthesis weights. This allows the model to leverage cross-sectional dependencies and shared information across multiple time series simultaneously. We develop an efficient Markov chain Monte Carlo (MCMC) algorithm for posterior inference, utilizing data augmentation and forward-filtering backward-sampling. Empirical applications to US inflation and global GDP growth demonstrate the improved performance of the proposed methods for quantile forecasting. In particular, FDRQS exhibits superior resilience during periods of extreme economic stress, such as the COVID-19 pandemic, by adaptively rebalancing agent contributions and capturing emergent global dependencies.

stat.ME

Contrastive Bayesian Inference for Unnormalized Models

Unnormalized (or energy-based) models provide a flexible framework for capturing the characteristics of data with complex dependency structures. However, the application of standard Bayesian inference methods has been severely limited because the parameter-dependent normalizing constant is either analytically intractable or computationally prohibitive to evaluate. A promising approach is score-based generalized Bayesian inference, which avoids evaluating the normalizing constant by replacing the likelihood with a scoring rule. However, this approach still requires careful tuning of the likelihood information, and it may fail to yield valid inference without appropriate control. To overcome this difficulty, we propose a fully Bayesian framework for inference on unnormalized models that does not require such tuning. We build on noise-contrastive estimation, which recasts inference as a binary classification problem between observed and noise samples, and treat the normalizing constant as an additional unknown parameter within the resulting likelihood. For exponential families, the classification likelihood becomes conditionally Gaussian via P\'olya-Gamma data augmentation, leading to a simple Gibbs sampler. We further establish posterior concentration and a Bernstein-von Mises theorem for our proposed method, providing theoretical justification for its uncertainty quantification. We demonstrate the proposed approach through two models: time-varying density models of temporal point processes and sparse torus graph models of multivariate circular data. Through simulation studies and real-data analyses, our proposed method provides accurate point estimates and principled uncertainty quantification.

stat.ME

Direct Bayesian Additive Regression Trees for Conditional Average Treatment Effects in Regression Discontinuity Designs

Regression discontinuity designs (RDD) are widely used for causal inference. In many empirical applications, treatment effects vary substantially with covariates, and ignoring such heterogeneity can lead to misleading conclusions, which motivates flexible modeling of heterogeneous treatment effects in RDD. To this end, we propose a Bayesian nonparametric approach to estimating heterogeneous treatment effects based on Bayesian Additive Regression Trees (BART). The key feature of our method lies in adopting a general Bayesian framework using a pseudo-model defined through a loss function for fitting local linear models around the cutoff, which gives direct modeling of heterogeneous treatment effects by BART. Optimal selection of the bandwidth parameter for the local model is implemented using the Hyv\"arinen score. Through numerical experiments, we demonstrate that the proposed approach flexibly captures complicated structures of heterogeneous treatment effects as a function of covariates.

stat.ME

Tree-Embedded Bayesian Factor Models for Multidimensional Categorical Distributions

Analyzing data collected from multiple observational units to estimate common and heterogeneous structures through a hierarchical model is a central task in Bayesian inference, and to this end, Bayesian factor models are one of the most widely used tools for this purpose. In this paper, we propose a novel Bayesian latent factor model for categorical distributions from grouped data, providing a parsimonious model for describing many observed distributions through lower-dimensional structures. Grouped data arise in a wide range of applications in social science, for example, distributions of age composition and income observed across locations. In these contexts, standard mixture models can be inefficient because the distributions do not necessarily exhibit clear clustering structures, and the distributions can be more accurately approximated as a combination of lower-dimensional characteristics. To analyze distribution-valued data with the Bayesian factor analysis, we adopt a tree-based transformation that embeds distributions into a Euclidean space and construct a Bayesian latent factor model in the transformed space. We develop the hierarchical model by incorporating the infinite factor model, which can adaptively estimate the number of effective factors. In addition, we propose its generalization by incorporating a spatial dependence by introducing a prior based on a SAR model. The proposed model provides smooth estimates of multivariate distributional structures, because once a tree-based transformation is applied, both univariate and multivariate distributions are essentially treated as the same Euclidean vectors. Through numerical experiments using real population data, we demonstrate that the proposed model outperforms existing parametric and Bayesian nonparametric models in various scenarios involving smooth spatial variations, especially under small sample sizes.

stat.ME

Bayesian Estimation of Variance under Fine Stratification via Mean-Variance Smoothing

Fine stratification survey is useful in many applications as its point estimator is unbiased, but the variance estimator under the design cannot be easily obtained, particularly when the sample size per stratum is as small as one unit. One common practice to overcome this difficulty is to collapse strata in pairs to create pseudo-strata and then estimate the variance. The estimator of variance achieved is not design-unbiased, and the positive bias increases as the population means of the paired pseudo-strata become more variant. The resulting confidence intervals can be unnecessarily large. In this paper, we propose a new Bayesian estimator for variance which does not rely on collapsing strata, unlike the previous methods given in the literature. We employ the penalized spline method for smoothing the mean and variance together in a nonparametric way. Furthermore, we make comparisons with the earlier work of Breidt et al. (2016). Throughout multiple simulation studies and an illustration using data from the National Survey of Family Growth (NSFG), we demonstrate the favorable performance of our methodology.

stat.ME

Difference-in-Differences under Local Dependence on Networks

Estimating causal effects under interference, where the stable unit treatment value assumption is violated, is critical in fields such as regional and public economics. Much of the existing research on causal inference under interference relies on a pre-specified "exposure mapping". This paper focuses on difference-in-difference and proposes a nonparametric identification strategy for direct and indirect average treatment effects under local interference on an observed network. In particular, we proposed a new concept of an indirect effect measuring the total outward influence of the intervension. Based on parallel trends assumption conditional on the neighborhood treatment vector, we develop inverse probability weighted and doubly robust estimators. We establish their asymptotic properties, including consistency under misspecification of nuisance models under some regularity conditions. Simulation studies and an empirical application demonstrate the effectiveness of the proposed method.

stat.ME

The Covariate-Assisted Bayesian Intransitive Bradley-Terry Model via Combinatorial Hodge Theory

Pairwise comparison data are widely used to recover latent rankings, yet the models in dominant use assume stochastic transitivity. When preferences are in fact intransitive, a single scalar strength conflates genuine hierarchy with cycle-induced structure, biasing both the recovered ranking and any covariate effects attributed to it. To address this limitation, we propose the Covariate-Assisted Bayesian Intransitive Bradley-Terry (CA-BIBT) model, which uses a combinatorial Hodge decomposition to resolve the latent match-up into identifiable and mutually orthogonal flows, attributing the component lying in the covariate-induced subspace to observed covariates and assigning the remaining components to the residuals. A global-local shrinkage prior on the residual cycle-induced flow adapts the model from transitive to intransitive regimes without prespecifying the regime, and a Gibbs sampler yields, as posterior byproducts, calibrated uncertainty for each flow, the posterior probability of the level at which the entities are rankable, and two complementary decision summaries that remain well-defined under intransitivity. In simulations, the CA-BIBT model recovers all flow components accurately with near-nominal coverage, and in applications to two animal dominance datasets, it distinguishes covariate-induced from residual cyclic dominance while quantifying posterior uncertainty, and demonstrates the practical utility of two complementary decision summaries.

stat.ME

Propensity Patchwork Kriging for Scalable Inference on Heterogeneous Treatment Effects

Gaussian process-based models are attractive for estimating heterogeneous treatment effects (HTE), but their computational cost limits scalability in causal inference settings. In this work, we address this challenge by extending Patchwork Kriging into the causal inference framework. Our proposed method partitions the data according to the estimated propensity score and applies Patchwork Kriging to enforce continuity of HTE estimates across adjacent regions. By imposing continuity constraints only along the propensity score dimension, rather than the full covariate space, the proposed approach substantially reduces computational cost while avoiding discontinuities inherent in simple local approximations. The resulting method can be interpreted as a smoothing extension of stratification and provides an efficient approach to HTE estimation. The proposed method is demonstrated through simulation studies and a real data application.

stat.ME

Repulsive g-Priors for Regression Mixtures

Mixture regression models are powerful tools for capturing heterogeneous covariate-response relationships, yet classical finite mixtures and Bayesian nonparametric alternatives often suffer from instability or overestimation of clusters when component separability is weak. Recent repulsive priors improve parsimony in density mixtures by discouraging nearby components, but their direct extension to regression is nontrivial since separation must respect the predictive geometry induced by covariates. We propose a repulsive g-prior for regression mixtures that enforces separation in the Mahalanobis metric, penalizing components indistinguishable in the predictive mean space. This construction preserves conjugacy-like updates while introducing geometry-aware interactions, enabling efficient blocked-collapsed Gibbs sampling. Theoretically, we establish tractable normalizing bounds, posterior contraction rates, and shrinkage of tail mass on the number of components. Simulations under correlated and overlapping designs demonstrate improved clustering and prediction relative to independent, Euclidean-repulsive, and sparsity-inducing baselines.

stat.ME

Robust Global Fr'echet Regression via Weight Regularization

The Fr\'echet regression is a useful method for modeling random objects in a general metric space given Euclidean covariates. However, the conventional approach could be sensitive to outlying objects in the sense that the distance from the regression surface is large compared to the other objects. In this study, we develop a robust version of the global Fr\'echet regression by incorporating weight parameters into the objective function. We then introduce the Elastic net regularization, favoring a sparse vector of robust parameters to control the influence of outlying objects. We provide a computational algorithm to iteratively estimate the regression function and weight parameters, with providing a linear convergence property. We also propose the Bayesian information criterion to select the tuning parameters for regularization, which gives adaptive robustness along with observed data. The finite sample performance of the proposed method is demonstrated through numerical studies on matrix and distribution responses.

stat.CO

On Misspecified Error Distributions in Bayesian Functional Clustering: Consequences and Remedies

Nonparametric Bayesian approaches provide a flexible framework for clustering without pre-specifying the number of groups, yet they are well known to overestimate the number of clusters, especially for functional data. We show that a fundamental cause of this phenomenon lies in misspecification of the error structure: errors are conventionally assumed to be independent across observed points in Bayesian functional models. Through high-dimensional clustering theory, we demonstrate that ignoring the underlying correlation leads to excess clusters regardless of the flexibility of prior distributions. Guided by this theory, we propose incorporating the underlying correlation structures via Gaussian processes and also present its scalable approximation with principled hyperparameter selection. Numerical experiments illustrate that even simple clustering based on Dirichlet processes performs well once error dependence is properly modeled.

stat.ME