SearcharxivSearch

arXiv subjects

Malay Ghosh

Publications and source records attributed to Malay Ghosh.

At least 19 recordsLinked to original sources

Asymptotic Minimax Estimation under Global-Local Priors

Global-local priors, often also referred to as shrinkage priors, have proved to be a very effective tool for the analysis of high dimensional data under sparsity. Asymptotic theoretical properties of such priors, studied under various scenarios, are now available in the literature. However, to our knowledge, theoretical guarantees of such priors provided so far, involve the assumption of known sample variance. The present paper relaxes this assumption, and carries out the analysis with a prior assigned to the error variance as well. In the process, some new tail bounds for shrinkage factors are developed, and these results are then utilized in providing asymptotic minimax rates for the posterior means of the parameters of interest.

math.ST

Conformity-Based Bayesian Projective Prediction

We propose a general robust prediction framework, termed conformity-based projective prediction (CPP), that integrates Bayesian predictive modeling with ideas from conformity-based conformal prediction. Rather than assessing conformity through residual-based scores, the CPP criterion defines conformity distributionally: a candidate value for a future response is considered conforming to the extent that its inclusion in the data leaves the leave-one-out predictive distributions of the observed responses undisturbed. The framework requires only that the leave-one-out and swapped predictive distributions are available in closed form and that the swapped predictive mean is differentiable in the candidate value. Under these conditions, we establish a general bounded-influence proposition and a general local convexity lemma, and prove that CPP dominates any plug-in predictor with unbounded influence in asymptotic variance under $\epsilon$-contamination models. When the posterior mean is linear in the observations -- as in Gaussian linear models, basis-expansion regression, and Gaussian process regression -- the swapped predictive mean is affine in the candidate value, yielding closed-form or one-dimensional optimization solutions and an efficient rank-two computational update; all general theoretical results specialize to explicit corollaries in this setting. Simulation experiments and two data analyses under the Gaussian linear model illustrate the finite-sample advantages of the proposed method, confirming the theoretical predictions across contamination levels, sample sizes, and predictor dimensions.

stat.ME

High-Dimensional Bernstein Von-Mises Theorems for Covariance and Precision Matrices

This paper aims to examine the characteristics of the posterior distribution of covariance/precision matrices in a "large $p$, large $n$" scenario, where $p$ represents the number of variables and $n$ is the sample size. Our analysis focuses on establishing asymptotic normality of the posterior distribution of the entire covariance/precision matrices under specific growth restrictions on $p_n$ and other mild assumptions. In particular, the limiting distribution turns out to be a symmetric matrix variate normal distribution whose parameters depend on the maximum likelihood estimate. Our results hold for a wide class of prior distributions which includes standard choices used by practitioners. Next, we consider Gaussian graphical models which induce sparsity in the precision matrix. Asymptotic normality of the corresponding posterior distribution is established under mild assumptions on the prior and true data-generating mechanism.

math.ST

Bias Corrected Variance Stabilizing Transformation for Small Area Estimation

Small area estimation models are typically based on the normality assumption of response variables. More recently, attention has been drawn to the transformation of the original variables to justify the assumption of normality. Variance stabilizing transformation of observation serves the dual purpose of reaching closer to normality, as well as known variance of the transformed variables in contrast to the assumption of known variances of the original variables, the latter needed to avoid non-identifiability. However, the existing literature on the topic ignores a certain bias introduced in the seemingly correct back transformation. The present paper rectifies this deficiency by introducing asymptotically unbiased empirical Bayes (EB) estimators of small area means. Mean squared errors (MSEs) and estimated MSEs of such estimators are provided. The theoretical results were accompanied with simulations and data analysis. A somewhat surprising phenomenon is a finding which connects one of our results to the natural exponential family quadratic variance function (NEF-QVF) family of distributions introduced by Morris (1982,1983).

stat.ME

Statistical Modeling of Combinatorial Response Data

There is a rich literature for modeling binary and polychotomous responses. However, existing methods are inadequate for handling combinatorial responses, where each response is an integer array under additional constraints. Such data are increasingly common in modern applications, such as surveys collected under skip logic, event propagation on a network, and observed matching in ecology. Ignoring the combinatorial structure leads to biased estimation and prediction. The fundamental challenge is the lack of a link function that connects a linear or functional predictor with a probability respecting the combinatorial constraints. In this article, we propose a novel augmented likelihood that views combinatorial response as a deterministic transform of a continuous latent variable. We specify the transform as the maximizer of integer linear program, and characterize useful properties such as dual thresholding representation. When taking a Bayesian approach and considering a multivariate normal distribution for the latent variable, our method becomes a direct generalization to the celebrated probit data augmentation, and enjoys straightforward computation via Markov chain Monte Carlo. We provide theoretical justification, including consistency and applicability, at an interesting intersection between duality and probability. We demonstrate the effectiveness of our method through simulations and a data application on the seasonal matching between waterfowl.

stat.ME

Posterior consistency in multi-response regression models with non-informative priors for the error covariance matrix in growing dimensions

The Inverse-Wishart (IW) distribution is a standard and popular choice of priors for covariance matrices and has attractive properties such as conditional conjugacy. However, the IW family of priors has crucial drawbacks, including the lack of effective choices for non-informative priors. Several classes of priors for covariance matrices that alleviate these drawbacks, while preserving computational tractability, have been proposed in the literature. These priors can be obtained through appropriate scale mixtures of IW priors. However, in the era of increasing dimensionality, the posterior consistency of models that incorporate such priors has not been investigated. We address this issue for the multi-response regression setting ($q$ responses, $n$ samples) under a wide variety of IW scale mixture priors for the error covariance matrix. Posterior consistency and contraction rates for both the regression coefficient matrix and the error covariance matrix are established in the ``large $q$, large $n$'' setting under mild assumptions on the true data-generating covariance matrix and relevant hyperparameters. In particular, the number of responses $q_n$ is allowed to grow with $n$, but with $q_n = o(n)$. Also, some results related to the inconsistency of the posterior distribution and posterior mean for $q_n/n \to γ$, where $γ\in (0,\infty)$ are provided.

math.ST

Bernstein von-Mises Theorem for g-prior and nonlocal prior

The paper develops Bernstein von Mises Theorem under hierarchical $g$ -priors for linear regression models. The results are obtained both when the error variance is known, and also when it is unknown. An inverse gamma prior is attached to the error variance in the later case. The paper also demonstrates some connection between the total variation and $α$-divergence measures.

math.ST

Global-Local Shrinkage Priors for Asymptotic Point and Interval Estimation of Normal Means under Sparsity

The paper addresses asymptotic estimation of normal means under sparsity. The primary focus is estimation of multivariate normal means where we obtain exact asymptotic minimax error under global-local shrinkage prior. This extends the corresponding univariate work of Ghosh and Chakrabarti (2017). In addition, we obtain similar results for the Dirichlet-Laplace prior as considered in Bhattacharya, Pati, Pillai, and Dunson (2015). Also, following van der Pas, Szabo, and van der Vaart (2017), we have been able to derive credible sets for multivariate normal means under global-local priors.

math.ST

Small Area Estimation under Square Root Transformed Fay-Herriot model with Functional Measurement Error in Covariates

We consider a small area estimation model under square-root transformation in the presence of functional measurement error. When measurement error is present, the Bayes predictor can no longer be used as it depends on the covariates even if parameters are known. Therefore suitable replacements are called for, and we propose a predictor that only depends on observed responses and data obtained from a large secondary survey. Moreover, some estimating methods of unknown parameters are considered. In the simulations section, We evaluate the performance using the mean squared prediction error (MSPE) and discuss several scenarios in terms of the number of areas and the sample size in a large secondary survey.

stat.ME

An Unbiased Predictor for Skewed Response Variable with Measurement Error in Covariate

We introduce a new small area predictor when the Fay-Herriot normal error model is fitted to a logarithmically transformed response variable, and the covariate is measured with error. This framework has been previously studied by Mosaferi et al. (2023). The empirical predictor given in their manuscript cannot perform uniformly better than the direct estimator. Our proposed predictor in this manuscript is unbiased and can perform uniformly better than the one proposed in Mosaferi et al. (2023). We derive an approximation of the mean squared error (MSE) for the predictor. The prediction intervals based on the MSE suffer from coverage problems. Thus, we propose a non-parametric bootstrap prediction interval which is more accurate. This problem is of great interest in small area applications since statistical agencies and agricultural surveys are often asked to produce estimates of right skewed variables with covariates measured with errors. With Monte Carlo simulation studies and two Census Bureau's data sets, we demonstrate the superiority of our proposed methodology.

stat.ME

High-dimensional properties for empirical priors in linear regression with unknown error variance

We study full Bayesian procedures for high-dimensional linear regression. We adopt data-dependent empirical priors introduced in [1]. In their paper, these priors have nice posterior contraction properties and are easy to compute. Our paper extend their theoretical results to the case of unknown error variance . Under proper sparsity assumption, we achieve model selection consistency, posterior contraction rates as well as Bernstein von-Mises theorem by analyzing multivariate t-distribution.

math.ST

Posterior Consistency for Bayesian Relevance Vector Machines

Statistical modeling and inference problems with sample sizes substantially smaller than the number of available covariates are challenging. Chakraborty et al. (2012) did a full hierarchical Bayesian analysis of nonlinear regression in such situations using relevance vector machines based on reproducing kernel Hilbert space (RKHS). But they did not provide any theoretical properties associated with their procedure. The present paper revisits their problem, introduces a new class of global-local priors different from theirs, and provides results on posterior consistency as well as posterior contraction rates

stat.ML

Contraction of a quasi-Bayesian model with shrinkage priors in precision matrix estimation

Currently several Bayesian approaches are available to estimate large sparse precision matrices, including Bayesian graphical Lasso (Wang, 2012), Bayesian structure learning (Banerjee and Ghosal, 2015), and graphical horseshoe (Li et al., 2019). Although these methods have exhibited nice empirical performances, in general they are computationally expensive. Moreover, we have limited knowledge about the theoretical properties, e.g., posterior contraction rate, of graphical Bayesian Lasso and graphical horseshoe. In this paper, we propose a new method that integrates some commonly used continuous shrinkage priors into a quasi-Bayesian framework featured by a pseudo-likelihood. Under mild conditions, we establish an optimal posterior contraction rate for the proposed method. Compared to existing approaches, our method has two main advantages. First, our method is computationally more efficient while achieving similar error rate; second, our framework is more amenable to theoretical analysis. Extensive simulation experiments and the analysis on a real data set are supportive of our theoretical results.

stat.ME

Transformed Fay-Herriot Model with Measurement Error in Covariates

Statistical agencies are often asked to produce small area estimates (SAEs) for positively skewed variables. When domain sample sizes are too small to support direct estimators, effects of skewness of the response variable can be large. As such, it is important to appropriately account for the distribution of the response variable given available auxiliary information. Motivated by this issue and in order to stabilize the skewness and achieve normality in the response variable, we propose an area-level log-measurement error model on the response variable. Then, under our proposed modeling framework, we derive an empirical Bayes (EB) predictor of positive small area quantities subject to the covariates containing measurement error. We propose a corresponding mean squared prediction error (MSPE) of EB predictor using both a jackknife and a bootstrap method. We show that the order of the bias is $O(m^{-1})$, where $m$ is the number of small areas. Finally, we investigate the performance of our methodology using both design-based and model-based simulation studies.

stat.ME

Ultra High-dimensional Multivariate Posterior Contraction Rate Under Shrinkage Priors

In recent years, shrinkage priors have received much attention in high-dimensional data analysis from a Bayesian perspective. Compared with widely used spike-and-slab priors, shrinkage priors have better computational efficiency. But the theoretical properties, especially posterior contraction rate, which is important in uncertainty quantification, are not established in many cases. In this paper, we apply global-local shrinkage priors to high-dimensional multivariate linear regression with unknown covariance matrix. We show that when the prior is highly concentrated near zero and has heavy tail, the posterior contraction rates for both coefficients matrix and covariance matrix are nearly optimal. Our results hold when number of features p grows much faster than the sample size n, which is of great interest in modern data analysis. We show that a class of readily implementable scale mixture of normal priors satisfies the conditions of the main theorem.

math.ST

Large-Scale Multiple Hypothesis Testing with the Normal-Beta Prime Prior

We revisit the problem of simultaneously testing the means of $n$ independent normal observations under sparsity. We take a Bayesian approach to this problem by introducing a scale-mixture prior known as the normal-beta prime (NBP) prior. We first derive new concentration properties when the beta prime density is employed for a scale parameter in Bayesian hierarchical models. To detect signals in our data, we then propose a hypothesis test based on thresholding the posterior shrinkage weight under the NBP prior. Taking the loss function to be the expected number of misclassified tests, we show that our test procedure asymptotically attains the optimal Bayes risk when the signal proportion $p$ is known. When $p$ is unknown, we introduce an empirical Bayes variant of our test which also asymptotically attains the Bayes Oracle risk in the entire range of sparsity parameters $p \propto n^{-ε}, ε\in (0, 1)$. Finally, we also consider restricted marginal maximum likelihood (REML) and hierarchical Bayes approaches for estimating a key hyperparameter in the NBP prior and examine multiple testing under these frameworks.

stat.ME

Consistent Bayesian Sparsity Selection for High-dimensional Gaussian DAG Models with Multiplicative and Beta-mixture Priors

Estimation of the covariance matrix for high-dimensional multivariate datasets is a challenging and important problem in modern statistics. In this paper, we focus on high-dimensional Gaussian DAG models where sparsity is induced on the Cholesky factor L of the inverse covariance matrix. In recent work, ([Cao, Khare, and Ghosh, 2019]), we established high-dimensional sparsity selection consistency for a hierarchical Bayesian DAG model, where an Erdos-Renyi prior is placed on the sparsity pattern in the Cholesky factor L, and a DAG-Wishart prior is placed on the resulting non-zero Cholesky entries. In this paper we significantly improve and extend this work, by (a) considering more diverse and effective priors on the sparsity pattern in L, namely the beta-mixture prior and the multiplicative prior, and (b) establishing sparsity selection consistency under significantly relaxed conditions on p, and the sparsity pattern of the true model. We demonstrate the validity of our theoretical results via numerical simulations, and also use further simulations to demonstrate that our sparsity selection approach is competitive with existing state-of-the-art methods including both frequentist and Bayesian approaches in various settings.

math.ST

High-dimensional posterior consistency for hierarchical non-local priors in regression

The choice of tuning parameters in Bayesian variable selection is a critical problem in modern statistics. In particular, for Bayesian linear regression with non-local priors, the scale parameter in the non-local prior density is an important tuning parameter which reflects the dispersion of the non-local prior density around zero, and implicitly determines the size of the regression coefficients that will be shrunk to zero. Current approaches treat the scale parameter as given, and suggest choices based on prior coverage/asymptotic considerations. In this paper, we consider the fully Bayesian approach introduced in (Wu, 2016) with the pMOM non-local prior and an appropriate Inverse-Gamma prior on the tuning parameter to analyze the underlying theoretical property. Under standard regularity assumptions, we establish strong model selection consistency in a high-dimensional setting, where $p$ is allowed to increase at a polynomial rate with n$or even at a sub-exponential rate with n. Through simulation studies, we demonstrate that our model selection procedure can outperform other Bayesian methods which treat the scale parameter as given, and commonly used penalized likelihood methods, in a range of simulation settings.

math.ST