SearcharxivSearch

arXiv subjects

Genya Kobayashi

Publications and source records attributed to Genya Kobayashi.

18 recordsLinked to original sources

Log-regularly varying scale mixture of asymmetric Laplaces for robust Bayesian quantile regression

Bayesian quantile regression based on the asymmetric Laplace (AL) distribution can be sensitive to extreme observations because of its exponentially decaying tails. We propose a robust error distribution constructed as a finite mixture of the AL distribution and a log-Pareto scale mixture of asymmetric Laplace distributions (LPAL). Unlike a direct log-Pareto extension of the normal location-scale representation of the AL distribution, the proposed AL-LPAL mixture preserves the prescribed quantile and exhibits log-regularly varying behavior in both tails. The LPAL component also has an unbounded density at the target quantile, yielding a distribution that combines sharp central concentration with super-heavy tails. We establish posterior robustness under arbitrarily extreme contamination and provide sufficient conditions for the existence of posterior moments of the regression coefficients and scale parameter. For posterior computation, we develop a Gibbs sampler using latent-variable augmentations and a computationally efficient mean-field variational Bayes approximation. Simulation studies show that the proposed method is competitive under moderate contamination and maintains stable point estimation with comparatively concentrated posterior intervals, particularly when severe contamination affects the quantile of interest. Applications to carbon dioxide and Boston housing data, using the same preprocessing as existing robust Bayesian quantile regression analyses, show favorable predictive performance across nearly all quantile levels and loss criteria considered.

stat.ME

Dynamic Bayesian regression quantile synthesis for forecasting outlook-at-risk

This paper proposes dynamic Bayesian regression quantile synthesis (DRQS), a novel method for quantile forecasting within the Bayesian predictive synthesis (BPS) framework designed to combine quantile-specific information from multiple agent models. While existing BPS approaches primarily focus on mean forecasting, our method directly targets the conditional quantiles of the response variable by utilizing the asymmetric Laplace distribution for the synthesis function. The resulting framework can be interpreted as a dynamic quantile linear model with latent predictors. We extend the univariate DRQS to a multivariate setting-factor DRQS (FDRQS)-by introducing a time-varying latent factor structure for the synthesis weights. This allows the model to leverage cross-sectional dependencies and shared information across multiple time series simultaneously. We develop an efficient Markov chain Monte Carlo (MCMC) algorithm for posterior inference, utilizing data augmentation and forward-filtering backward-sampling. Empirical applications to US inflation and global GDP growth demonstrate the improved performance of the proposed methods for quantile forecasting. In particular, FDRQS exhibits superior resilience during periods of extreme economic stress, such as the COVID-19 pandemic, by adaptively rebalancing agent contributions and capturing emergent global dependencies.

stat.ME

Tree-Embedded Bayesian Factor Models for Multidimensional Categorical Distributions

Analyzing data collected from multiple observational units to estimate common and heterogeneous structures through a hierarchical model is a central task in Bayesian inference, and to this end, Bayesian factor models are one of the most widely used tools for this purpose. In this paper, we propose a novel Bayesian latent factor model for categorical distributions from grouped data, providing a parsimonious model for describing many observed distributions through lower-dimensional structures. Grouped data arise in a wide range of applications in social science, for example, distributions of age composition and income observed across locations. In these contexts, standard mixture models can be inefficient because the distributions do not necessarily exhibit clear clustering structures, and the distributions can be more accurately approximated as a combination of lower-dimensional characteristics. To analyze distribution-valued data with the Bayesian factor analysis, we adopt a tree-based transformation that embeds distributions into a Euclidean space and construct a Bayesian latent factor model in the transformed space. We develop the hierarchical model by incorporating the infinite factor model, which can adaptively estimate the number of effective factors. In addition, we propose its generalization by incorporating a spatial dependence by introducing a prior based on a SAR model. The proposed model provides smooth estimates of multivariate distributional structures, because once a tree-based transformation is applied, both univariate and multivariate distributions are essentially treated as the same Euclidean vectors. Through numerical experiments using real population data, we demonstrate that the proposed model outperforms existing parametric and Bayesian nonparametric models in various scenarios involving smooth spatial variations, especially under small sample sizes.

stat.ME

General Bayesian quantile regression for counts via generative modeling

Count data frequently arises in biomedical applications, such as the length of hospital stay. However, their discrete nature poses significant challenges for appropriately modeling conditional quantiles, which are crucial for understanding heterogeneous effects and variability in outcomes. To solve the practical difficulty, we propose a novel general Bayesian framework for quantile regression tailored to count data. We seek the regression parameter on the conditional quantile by minimizing the expected loss with respect to the distribution of the conditional quantile of the latent continuous variable associated with the observed count response variable. By modeling the unknown conditional distribution through a Bayesian nonparametric kernel mixture for the joint distribution of the count response and covariates, we obtain the posterior distribution of the regression parameter via a simple optimization. We numerically demonstrate that the proposed method improves bias and estimation accuracy of the existing crude approaches to count quantile regression. Furthermore, we analyze the length of hospital stay for acute myocardial infarction and demonstrate that the proposed method gives more interpretable and flexible results than the existing ones.

stat.ME

Predicting COVID-19 hospitalisation using a mixture of Bayesian predictive syntheses

This paper proposes a novel methodology called the mixture of Bayesian predictive syntheses (MBPS) for multiple time series count data for the challenging task of predicting the numbers of COVID-19 inpatients and isolated cases in Japan and Korea at the subnational-level. MBPS combines a set of predictive models and partitions the multiple time series into clusters based on their contribution to predicting the outcome. In this way, MBPS leverages the shared information within each cluster and is suitable for predicting COVID-19 inpatients since the data exhibit similar dynamics over multiple areas. Also, MBPS avoids using a multivariate count model, which is generally cumbersome to develop and implement. Our Japanese and Korean data analyses demonstrate that the proposed MBPS methodology has improved predictive accuracy and uncertainty quantification.

stat.AP

Bayesian Benchmarking Small Area Estimation via Entropic Tilting

Benchmarking estimation and its risk evaluation is a practically important issue in small area estimation. While Bayesian methods have been widely adopted in small area estimation, existing benchmarking approaches are often ad-hoc, such as projecting each MCMC draw to satisfy the constraint. In contrast, our work provides a unified Bayesian formulation based on entropic tilting, which offers a more principled way to define the benchmarked posterior distribution. This approach yields benchmarked point estimates together with coherent uncertainty quantification. We first introduce general Monte Carlo methods for obtaining a benchmarked posterior under hierarchical Bayesian approaches and then show that the benchmarked posterior under empirical Bayesian frameworks can be obtained in an analytical form for some small area models. We demonstrate the usefulness of the proposed method through simulation and empirical studies.

stat.ME

Bayesian factor zero-inflated Poisson model for multiple grouped count data

This paper proposes a computationally efficient Bayesian factor model for multiple grouped count data. Adopting the link function approach, the proposed model can capture the association within and between the at-risk probabilities and Poisson counts over multiple dimensions. The likelihood function for the grouped count data consists of the differences of the cumulative distribution functions evaluated at the endpoints of the groups, defining the probabilities of each data point falling in the groups. The combination of the data augmentation of underlying counts, the Pólya-Gamma augmentation to approximate the Poisson distribution, and parameter expansion for the factor components is used to facilitate posterior computing. The efficacy of the proposed factor model is demonstrated using the simulated data and real data on the involvement of youths in the nineteen illegal activities.

stat.ME

Similarity-based Random Partition Distribution for Clustering Functional Data

Random partition distribution is a crucial tool for model-based clustering. This study advances the field of random partition in the context of functional spatial data, focusing on the challenges posed by hourly population data across various regions and dates. We propose an extension of the generalized Dirichlet process, named the similarity-based generalized Dirichlet process (SGDP)-type distribution, to address the limitations of simple random partition distributions (e.g., those induced by the Dirichlet process), such as an overabundance of clusters. This model prevents excess cluster production and incorporates pairwise similarity information to ensure accurate and meaningful clustering. The theoretical properties of the SGDP-type distribution are studied. Then, SGDP-type random partition is applied to a real-world dataset of hourly population flow in $500\text{m}$ meshes in the central part of Tokyo. In this empirical context, our method excels at detecting meaningful patterns in the data while accounting for spatial nuances. The results underscore the adaptability and utility of the method, showcasing its prowess in revealing intricate spatiotemporal dynamics. The proposed random partition will significantly contribute to urban planning, transportation, and policy-making and will be a helpful tool for understanding population dynamics and their implications.

stat.ME

Spatio-temporal smoothing, interpolation and prediction of income distributions based on grouped data

The Housing and Land Survey (HLS) of Japan provides municipality-level grouped data on household incomes. Although such data can be invaluable for effective local policymaking, their analysis is often hindered by several challenges, including limited information inherent in the grouped format, the presence of missing areas, and the low frequency of survey implementation. To address these issues, we propose a novel grouped-data-based spatio-temporal finite mixture model for estimating income distributions across multiple spatial units and time points. A unique feature of the proposed method is that all areas share common latent distributions, while the mixing proportions, incorporating spatial and temporal effects, capture the potential area-wise heterogeneity. Consequently, the inclusion of these effects enables smoothing quantities of interest over space and time, imputing missing values, and predicting future trends. By applying the proposed method to the HLS data, we generate complete maps of income and inequality measures at any given time, thereby facilitating rapid and efficient policymaking with fine granularity.

stat.AP

Robust Fitting of Mixture Models using Weighted Complete Estimating Equations

Mixture modeling, which considers the potential heterogeneity in data, is widely adopted for classification and clustering problems. Mixture models can be estimated using the Expectation-Maximization algorithm, which works with the complete estimating equations conditioned by the latent membership variables of the cluster assignment based on the hierarchical expression of mixture models. However, when the mixture components have light tails such as a normal distribution, the mixture model can be sensitive to outliers. This study proposes a method of weighted complete estimating equations (WCE) for the robust fitting of mixture models. Our WCE introduces weights to complete estimating equations such that the weights can automatically downweight the outliers. The weights are constructed similarly to the density power divergence for mixture models, but in our WCE, they depend only on the component distributions and not on the whole mixture. A novel expectation-estimating-equation (EEE) algorithm is also developed to solve the WCE. For illustrative purposes, a multivariate Gaussian mixture, a mixture of experts, and a multivariate skew normal mixture are considered, and how our EEE algorithm can be implemented for these specific models is described. The numerical performance of the proposed robust estimation method was examined using simulated and real datasets.

stat.ME

Predicting Infection of COVID-19 in Japan: State Space Modeling Approach

The number of confirmed cases of the coronavirus disease (COVID-19) in Japan has been increasing day by day and has had a serious impact on the society especially after the declaration of the state of emergency on April 7, 2020. This study analyzes the real time data from March 1 to April 22, 2020 by adopting a sophisticated statistical modeling tool based on the state space model combined with the well-known susceptible-exposed-infected (SIR) model. The model estimation and forecasting are conducted using the Bayesian methodology. The present study provides the parameter estimates of the unknown parameters that critically determine the epidemic process derived from the SIR model and prediction of the future transition of the infectious proportion including the size and timing of the epidemic peak with the prediction intervals that naturally accounts for the uncertainty. The prediction results under various scenarios reveals that the temporary reduction in the infection rate until the planned lifting of the state on May 6 will only delay the epidemic peak slightly. In order to minimize the spread of the epidemic, it is strongly suggested that an intervention is carried out for an extended period of time and that the government and individuals make a long term effort to reduce the infection rate even after the lifting.

stat.AP

Small area estimation of general finite-population parameters based on grouped data

This paper proposes a new model-based approach to small area estimation of general finite-population parameters based on grouped data or frequency data, which is often available from sample surveys. Grouped data contains information on frequencies of some pre-specified groups in each area, for example the numbers of households in the income classes, and thus provides more detailed insight about small areas than area-level aggregated data. A direct application of the widely used small area methods, such as the Fay-Herriot model for area-level data and nested error regression model for unit-level data, is not appropriate since they are not designed for grouped data. The newly proposed method adopts the multinomial likelihood function for the grouped data. In order to connect the group probabilities of the multinomial likelihood and the auxiliary variables within the framework of small area estimation, we introduce the unobserved unit-level quantities of interest which follows the linear mixed model with the random intercepts and dispersions after some transformation. Then the probabilities that a unit belongs to the groups can be derived and are used to construct the likelihood function for the grouped data given the random effects. The unknown model parameters (hyperparameters) are estimated by a newly developed Monte Carlo EM algorithm using an efficient importance sampling. The empirical best predicts (empirical Bayes estimates) of small area parameters can be calculated by a simple Gibbs sampling algorithm. The numerical performance of the proposed method is illustrated based on the model-based and design-based simulations. In the application to the city level grouped income data of Japan, we complete the patchy maps of the Gini coefficient as well as mean income across the country.

stat.ME

Bayesian approach to Lorenz curve using time series grouped data

This study is concerned with estimating the inequality measures associated with the underlying hypothetical income distribution from the times series grouped data on the Lorenz curve. We adopt the Dirichlet pseudo likelihood approach where the parameters of the Dirichlet likelihood are set to the differences between the Lorenz curve of the hypothetical income distribution for the consecutive income classes and propose a state space model which combines the transformed parameters of the Lorenz curve through a time series structure. Furthermore, the information on the sample size in each survey is introduced into the originally nuisance Dirichlet precision parameter to take into account the variability from the sampling. From the simulated data and real data on the Japanese monthly income survey, it is confirmed that the proposed model produces more efficient estimates on the inequality measures than the existing models without time series structures.

stat.ME

Estimation and inference for area-wise spatial income distributions from grouped data

Estimating income distributions plays an important role in the measurement of inequality and poverty over space. The existing literature on income distributions predominantly focuses on estimating an income distribution for a country or a region separately and the simultaneous estimation of multiple income distributions has not been discussed in spite of its practical importance. In this work, we develop an effective method for the simultaneous estimation and inference for area-wise spatial income distributions taking account of geographical information from grouped data. Based on the multinomial likelihood function for grouped data, we propose a spatial state-space model for area-wise parameters of parametric income distributions. We provide an efficient Bayesian approach to estimation and inference for area-wise latent parameters, which enables us to compute area-wise summary measures of income distributions such as mean incomes and Gini indices, not only for sampled areas but also for areas without any samples thanks to the latent spatial state-space structure. The proposed method is demonstrated using the Japanese municipality-wise grouped income data. The simulation studies show the superiority of the proposed method to a crude conventional approach which estimates the income distributions separately.

stat.ME

Approximate Bayesian Computation for Lorenz Curves from Grouped Data

This paper proposes a new Bayesian approach to estimate the Gini coefficient from the Lorenz curve based on grouped data. The proposed approach assumes a hypothetical income distribution and estimates the parameter by directly working on the likelihood function implied by the Lorenz curve of the income distribution from the grouped data. It inherits the advantages of two existing approaches through which the Gini coefficient can be estimated more accurately and a straightforward interpretation about the underlying income distribution is provided. Since the likelihood function is implicitly defined, the approximate Bayesian computational approach based on the sequential Monte Carlo method is adopted. The usefulness of the proposed approach is illustrated through the simulation study and the Japanese income data.

stat.AP

Latent Mixture Modeling for Clustered Data

This article proposes a mixture modeling approach to estimating cluster-wise conditional distributions in clustered (grouped) data. We adapt the mixture-of-experts model to the latent distributions, and propose a model in which each cluster-wise density is represented as a mixture of latent experts with cluster-wise mixing proportions distributed as Dirichlet distribution. The model parameters are estimated by maximizing the marginal likelihood function using a newly developed Monte Carlo Expectation-Maximization algorithm. We also extend the model such that the distribution of cluster-wise mixing proportions depends on some cluster-level covariates. The finite sample performance of the proposed model is compared with some existing mixture modeling approaches as well as linear mixed model through the simulation studies. The proposed model is also illustrated with the posted land price data in Japan.

stat.ME

Bayesian Nonparametric Instrumental Variable Regression Approach to Quantile Inference

This study extends the Bayesian nonparametric instrumental variable regression model to determine the structural effects of covariates on the conditional quantile of the response variable. The error distribution is nonparametrically modelled using a Dirichlet mixture of bivariate normal distributions. The mean functions include the smooth effects of the covariates represented using the spline functions in an additive manner. The conditional variance of the second-stage error is also modelled using the spline functions such that it varies smoothly with the covariates. Accordingly, the proposed model allows for considerable flexibility in the shape of the quantile function while correcting for an endogeneity effect. The posterior inference for the proposed model is based on the Markov chain Monte Carlo method that requires no Metropolis-Hastings update. The approach is demonstrated using simulated and real data on the death rate in Japan during the inter-war period.

stat.ME

Bayesian Endogenous Tobit Quantile Regression

This study proposes $p$-th Tobit quantile regression models with endogenous variables. In the first stage regression of the endogenous variable on the exogenous variables, the assumption that the $α$-th quantile of the error term is zero is introduced. Then, the residual of this regression model is included in the $p$-th quantile regression model in such a way that the $p$-th conditional quantile of the new error term is zero. The error distribution of the first stage regression is modelled around the zero $α$-th quantile assumption by using parametric and semiparametric approaches. Since the value of $α$ is a priori unknown, it is treated as an additional parameter and is estimated from the data. The proposed models are then demonstrated by using simulated data and real data on the labour supply of married women.

stat.ME