Searcharxiv⌕ Search

arXiv subjects

Yuki Kawakubo

Publications and source records attributed to Yuki Kawakubo.

11 recordsLinked to original sources

Spatio-temporal smoothing, interpolation and prediction of income distributions based on grouped data

The Housing and Land Survey (HLS) of Japan provides municipality-level grouped data on household incomes. Although such data can be invaluable for effective local policymaking, their analysis is often hindered by several challenges, including limited information inherent in the grouped format, the presence of missing areas, and the low frequency of survey implementation. To address these issues, we propose a novel grouped-data-based spatio-temporal finite mixture model for estimating income distributions across multiple spatial units and time points. A unique feature of the proposed method is that all areas share common latent distributions, while the mixing proportions, incorporating spatial and temporal effects, capture the potential area-wise heterogeneity. Consequently, the inclusion of these effects enables smoothing quantities of interest over space and time, imputing missing values, and predicting future trends. By applying the proposed method to the HLS data, we generate complete maps of income and inequality measures at any given time, thereby facilitating rapid and efficient policymaking with fine granularity.

stat.AP↗

Bayesian Benchmarking Small Area Estimation via Entropic Tilting

Benchmarking estimation and its risk evaluation is a practically important issue in small area estimation. While Bayesian methods have been widely adopted in small area estimation, existing benchmarking approaches are often ad-hoc, such as projecting each MCMC draw to satisfy the constraint. In contrast, our work provides a unified Bayesian formulation based on entropic tilting, which offers a more principled way to define the benchmarked posterior distribution. This approach yields benchmarked point estimates together with coherent uncertainty quantification. We first introduce general Monte Carlo methods for obtaining a benchmarked posterior under hierarchical Bayesian approaches and then show that the benchmarked posterior under empirical Bayesian frameworks can be obtained in an analytical form for some small area models. We demonstrate the usefulness of the proposed method through simulation and empirical studies.

stat.ME↗

Multilevel Decomposition of Generalized Entropy Measures Using Constrained Bayes Estimation: An Application to Japanese Regional Data

We propose a method for multilevel decomposition of generalized entropy (GE) measures that explicitly accounts for nested population structures such as national, regional, and subregional levels. Standard approaches that estimate GE separately at each level do not guarantee compatibility with multilevel decomposition. Our method constrains lower-level GE estimates to match higher-level benchmarks while preserving hierarchical relationships across layers. We apply the method to Japanese income data to estimate GE at the national, prefectural, and municipal levels, decomposing national inequality into between-prefecture and within-prefecture inequality, and further decomposing prefectural GE into between-municipality and within-municipality inequality.

econ.EM↗

Predicting COVID-19 hospitalisation using a mixture of Bayesian predictive syntheses

This paper proposes a novel methodology called the mixture of Bayesian predictive syntheses (MBPS) for multiple time series count data for the challenging task of predicting the numbers of COVID-19 inpatients and isolated cases in Japan and Korea at the subnational-level. MBPS combines a set of predictive models and partitions the multiple time series into clusters based on their contribution to predicting the outcome. In this way, MBPS leverages the shared information within each cluster and is suitable for predicting COVID-19 inpatients since the data exhibit similar dynamics over multiple areas. Also, MBPS avoids using a multivariate count model, which is generally cumbersome to develop and implement. Our Japanese and Korean data analyses demonstrate that the proposed MBPS methodology has improved predictive accuracy and uncertainty quantification.

stat.AP↗

Small Area Estimation with Spatially Varying Natural Exponential Families

Two-stage hierarchical models have been widely used in small area estimation to produce indirect estimates of areal means. When the areas are treated exchangeably and the model parameters are assumed to be the same over all areas, we might lose the efficiency in the presence of spatial heterogeneity. To overcome this problem, we consider a two-stage area-level model based on natural exponential family with spatially varying model parameters. We employ geographically weighted regression approach to estimating the varying parameters and suggest a new empirical Bayes estimator of the areal mean. We also discuss some related problems, including the mean squared error estimation, benchmarked estimation, and estimation in non-sampled areas. The performance of the proposed method is evaluated through simulations and applications to two data sets.

stat.ME↗

Small area estimation of general finite-population parameters based on grouped data

This paper proposes a new model-based approach to small area estimation of general finite-population parameters based on grouped data or frequency data, which is often available from sample surveys. Grouped data contains information on frequencies of some pre-specified groups in each area, for example the numbers of households in the income classes, and thus provides more detailed insight about small areas than area-level aggregated data. A direct application of the widely used small area methods, such as the Fay-Herriot model for area-level data and nested error regression model for unit-level data, is not appropriate since they are not designed for grouped data. The newly proposed method adopts the multinomial likelihood function for the grouped data. In order to connect the group probabilities of the multinomial likelihood and the auxiliary variables within the framework of small area estimation, we introduce the unobserved unit-level quantities of interest which follows the linear mixed model with the random intercepts and dispersions after some transformation. Then the probabilities that a unit belongs to the groups can be derived and are used to construct the likelihood function for the grouped data given the random effects. The unknown model parameters (hyperparameters) are estimated by a newly developed Monte Carlo EM algorithm using an efficient importance sampling. The empirical best predicts (empirical Bayes estimates) of small area parameters can be calculated by a simple Gibbs sampling algorithm. The numerical performance of the proposed method is illustrated based on the model-based and design-based simulations. In the application to the city level grouped income data of Japan, we complete the patchy maps of the Gini coefficient as well as mean income across the country.

stat.ME↗

Bayesian approach to Lorenz curve using time series grouped data

This study is concerned with estimating the inequality measures associated with the underlying hypothetical income distribution from the times series grouped data on the Lorenz curve. We adopt the Dirichlet pseudo likelihood approach where the parameters of the Dirichlet likelihood are set to the differences between the Lorenz curve of the hypothetical income distribution for the consecutive income classes and propose a state space model which combines the transformed parameters of the Lorenz curve through a time series structure. Furthermore, the information on the sample size in each survey is introduced into the originally nuisance Dirichlet precision parameter to take into account the variability from the sampling. From the simulated data and real data on the Japanese monthly income survey, it is confirmed that the proposed model produces more efficient estimates on the inequality measures than the existing models without time series structures.

stat.ME↗

Estimation and inference for area-wise spatial income distributions from grouped data

Estimating income distributions plays an important role in the measurement of inequality and poverty over space. The existing literature on income distributions predominantly focuses on estimating an income distribution for a country or a region separately and the simultaneous estimation of multiple income distributions has not been discussed in spite of its practical importance. In this work, we develop an effective method for the simultaneous estimation and inference for area-wise spatial income distributions taking account of geographical information from grouped data. Based on the multinomial likelihood function for grouped data, we propose a spatial state-space model for area-wise parameters of parametric income distributions. We provide an efficient Bayesian approach to estimation and inference for area-wise latent parameters, which enables us to compute area-wise summary measures of income distributions such as mean incomes and Gini indices, not only for sampled areas but also for areas without any samples thanks to the latent spatial state-space structure. The proposed method is demonstrated using the Japanese municipality-wise grouped income data. The simulation studies show the superiority of the proposed method to a crude conventional approach which estimates the income distributions separately.

stat.ME↗

Conditional Akaike information under covariate shift with application to small area estimation

In this study, we consider the problem of selecting explanatory variables of fixed effects in linear mixed models under covariate shift, which is when the values of covariates in the model for prediction differ from those in the model for observed data. We construct a variable selection criterion based on the conditional Akaike information introduced by Vaida and Blanchard (2005). We focus especially on covariate shift in small area estimation and demonstrate the usefulness of the proposed criterion. In addition, numerical performance is investigated through simulations, one of which is a design-based simulation using a real dataset of land prices.

stat.ME↗

A Variant of AIC based on the Bayesian Marginal Likelihood

We propose information criteria that measure the prediction risk of a predictive density based on the Bayesian marginal likelihood from a frequentist point of view. We derive criteria for selecting variables in linear regression models, assuming a prior distribution of the regression coefficients. Then, we discuss the relationship between the proposed criteria and related criteria. There are three advantages of our method. First, this is a compromise between the frequentist and Bayesian standpoints because it evaluates the frequentist's risk of the Bayesian model. Thus, it is less influenced by a prior misspecification. Second, the criteria exhibits consistency when selecting the true model. Third, when a uniform prior is assumed for the regression coefficients, the resulting criterion is equivalent to the residual information criterion (RIC) of Shi and Tsai (2002).

stat.ME↗

Latent Mixture Modeling for Clustered Data

This article proposes a mixture modeling approach to estimating cluster-wise conditional distributions in clustered (grouped) data. We adapt the mixture-of-experts model to the latent distributions, and propose a model in which each cluster-wise density is represented as a mixture of latent experts with cluster-wise mixing proportions distributed as Dirichlet distribution. The model parameters are estimated by maximizing the marginal likelihood function using a newly developed Monte Carlo Expectation-Maximization algorithm. We also extend the model such that the distribution of cluster-wise mixing proportions depends on some cluster-level covariates. The finite sample performance of the proposed model is compared with some existing mixture modeling approaches as well as linear mixed model through the simulation studies. The proposed model is also illustrated with the posted land price data in Japan.

stat.ME↗