SearcharxivSearch

arXiv subjects

Gauri Sankar Datta

Publications and source records attributed to Gauri Sankar Datta.

5 recordsLinked to original sources

Credible Distributions of Overall Ranking of Entities

Ranking, and inferences based on ranking of a set of entities, are important problems in numerous contexts. This is especially true in small area statistics where there may be only a limited amount of directly observed data from each entity or small area, while precise and accurate estimates of best or worst performing entities are needed for fund allocation, planning and policymaking, stakeholder advocacy, evaluation of welfare programs, and so on. However, ranks estimates constructed exclusively on point estimates of parameters lack uncertainty quantification, and may lead to imbalances and inequities when these are based on small sample sizes. We propose novel Bayesian approaches to address this problem. Our proposals result in partitions of the parameter space with posterior distribution driven partial ordering of the sets in a partition. This in turn translates to a coherent probability mass function over ranks for every entity, and a coherent probability mass function over entities for every rank. Our Bayesian algorithms significantly outperform the state-of-the-art non-Bayesian alternatives, and are amenable to inclusion of covariates in the model as well as borrowing strengths across small areas. We evaluate our proposed Bayesian algorithms in terms of accuracy and stability using a number of applications and a simulation study. Additionally, we develop a novel theoretical framework for inference and ranking problems involving a triangular array of Fay-Herriot models and data, and provide probabilistic guarantees of performances of the proposed Bayesian ranking algorithms.

stat.ME

A Hierarchical Bayes Unit-Level Small Area Estimation Model for Normal Mixture Populations

National statistical agencies are regularly required to produce estimates about various subpopulations, formed by demographic and/or geographic classifications, based on a limited number of samples. Traditional direct estimates computed using only sampled data from individual subpopulations are usually unreliable due to small sample sizes. Subpopulations with small samples are termed small areas or small domains. To improve on the less reliable direct estimates, model-based estimates, which borrow information from suitable auxiliary variables, have been extensively proposed in the literature. However, standard model-based estimates rely on the normality assumptions of the error terms. In this research we propose a hierarchical Bayesian (HB) method for the unit-level nested error regression model based on a normal mixture for the unit-level error distribution. To implement our proposal we use a uniform prior for the regression parameters, random effects variance parameter, and the mixing proportion, and we use a partially proper non-informative prior distribution for the two unit-level error variance components in the mixture. We apply our method to two examples to predict summary characteristics of farm products at the small area level. One of the examples is prediction of twelve county-level crop areas cultivated for corn in some Iowa counties. The other example involves total cash associated in farm operations in twenty-seven farming regions in Australia. We compare predictions of small area characteristics based on the proposed method with those obtained by applying the Datta and Ghosh (1991) and the Chakraborty et al. (2018) HB methods. Our simulation study comparing these three Bayesian methods showed the superiority of our proposed method, measured by prediction mean squared error, coverage probabilities and lengths of credible intervals for the small area means.

stat.ME

Robust Hierarchical Bayes Small Area Estimation for Nested Error Regression Model

National statistical institutes in many countries are now mandated to produce reliable statistics for important variables such as population, income, unemployment, health outcomes, etc. for small areas, defined by geography and/or demography. Due to small samples from these areas, direct sample-based estimates are often unreliable. Model-based small area estimation is now extensively used to generate reliable statistics by "borrowing strength" from other areas and related variables through suitable models. Outliers adversely influence standard model-based small area estimates. To deal with outliers, Sinha and Rao (2009) proposed a robust frequentist approach. In this article, we present a robust Bayesian alternative to the nested error regression model for unit-level data to mitigate outliers. We consider a two-component scale mixture of normal distributions for the unit-level error to model outliers and present a computational approach to produce Bayesian predictors of small area means under a noninformative prior for model parameters. A real example and extensive simulations convincingly show robustness of our Bayesian predictors to outliers. Simulations comparison of these two procedures with Bayesian predictors by Datta and Ghosh (1991) and M-quantile estimators by Chambers et al. (2014) shows that our proposed procedure is better than the others in terms of bias, variability, and coverage probability of prediction intervals, when there are outliers. The superior frequentist performance of our procedure shows its dual (Bayes and frequentist) dominance, and makes it attractive to all practitioners, both Bayesian and frequentist, of small area estimation.

stat.ME

A two-component normal mixture alternative to the Fay-Herriot model

This article considers a robust hierarchical Bayesian approach to deal with random effects of small area means when some of these effects assume extreme values, resulting in outliers. In presence of outliers, the standard Fay-Herriot model, used for modeling area-level data, under normality assumptions of the random effects may overestimate random effects variance, thus provides less than ideal shrinkage towards the synthetic regression predictions and inhibits borrowing information. Even a small number of substantive outliers of random effects result in a large estimate of the random effects variance in the Fay-Herriot model, thereby achieving little shrinkage to the synthetic part of the model or little reduction in posterior variance associated with the regular Bayes estimator for any of the small areas. While a scale mixture of normal distributions with known mixing distribution for the random effects has been found to be effective in presence of outliers, the solution depends on the mixing distribution. As a possible alternative solution to the problem, a two-component normal mixture model has been proposed based on noninformative priors on the model variance parameters, regression coefficients and the mixing probability. Data analysis and simulation studies based on real, simulated and synthetic data show advantage of the proposed method over the standard Bayesian Fay-Herriot solution derived under normality of random effects.

stat.ME