SearcharxivSearch

arXiv subjects

Hani Doss

Publications and source records attributed to Hani Doss.

4 recordsLinked to original sources

An Empirical Bayes Approach to Shrinkage Estimation on the Manifold of Symmetric Positive-Definite Matrices

The James-Stein estimator is an estimator of the multivariate normal mean and dominates the maximum likelihood estimator (MLE) under squared error loss. The original work inspired great interest in developing shrinkage estimators for a variety of problems. Nonetheless, research on shrinkage estimation for manifold-valued data is scarce. In this paper, we propose shrinkage estimators for the parameters of the Log-Normal distribution defined on the manifold of $N \times N$ symmetric positive-definite matrices. For this manifold, we choose the Log-Euclidean metric as its Riemannian metric since it is easy to compute and is widely used in applications. By using the Log-Euclidean distance in the loss function, we derive a shrinkage estimator in an analytic form and show that it is asymptotically optimal within a large class of estimators including the MLE, which is the sample Fréchet mean of the data. We demonstrate the performance of the proposed shrinkage estimator via several simulated data experiments. Furthermore, we apply the shrinkage estimator to perform statistical inference in diffusion magnetic resonance imaging problems.

math.ST

An MCMC Approach to Empirical Bayes Inference and Bayesian Sensitivity Analysis via Empirical Processes

Consider a Bayesian situation in which we observe $Y \sim p_θ$, where $θ\in Θ$, and we have a family $\{ ν_h, \, h \in \mathcal{H} \}$ of potential prior distributions on $Θ$. Let $g$ be a real-valued function of $θ$, and let $I_g(h)$ be the posterior expectation of $g(θ)$ when the prior is $ν_h$. We are interested in two problems: (i) selecting a particular value of $h$, and (ii) estimating the family of posterior expectations $\{ I_g(h), \, h \in \mathcal{H} \}$. Let $m_y(h)$ be the marginal likelihood of the hyperparameter $h$: $m_y(h) = \int p_θ(y) \, ν_h(dθ)$. The empirical Bayes estimate of $h$ is, by definition, the value of $h$ that maximizes $m_y(h)$. It turns out that it is typically possible to use Markov chain Monte Carlo to form point estimates for $m_y(h)$ and $I_g(h)$ for each individual $h$ in a continuum, and also confidence intervals for $m_y(h)$ and $I_g(h)$ that are valid pointwise. However, we are interested in forming estimates, with confidence statements, of the entire families of integrals $\{ m_y(h), \, h \in \mathcal{H} \}$ and $\{ I_g(h), \, h \in \mathcal{H} \}$: we need estimates of the first family in order to carry out empirical Bayes inference, and we need estimates of the second family in order to do Bayesian sensitivity analysis. We establish strong consistency and functional central limit theorems for estimates of these families by using tools from empirical process theory. We give two applications, one to Latent Dirichlet Allocation, which is used in topic modelling, and the other is to a model for Bayesian variable selection in linear regression.

stat.ME

Computational approaches for empirical Bayes methods and Bayesian sensitivity analysis

We consider situations in Bayesian analysis where we have a family of priors $ν_h$ on the parameter $θ$, where $h$ varies continuously over a space $\mathcal{H}$, and we deal with two related problems. The first involves sensitivity analysis and is stated as follows. Suppose we fix a function $f$ of $θ$. How do we efficiently estimate the posterior expectation of $f(θ)$ simultaneously for all $h$ in $\mathcal{H}$? The second problem is how do we identify subsets of $\mathcal{H}$ which give rise to reasonable choices of $ν_h$? We assume that we are able to generate Markov chain samples from the posterior for a finite number of the priors, and we develop a methodology, based on a combination of importance sampling and the use of control variates, for dealing with these two problems. The methodology applies very generally, and we show how it applies in particular to a commonly used model for variable selection in Bayesian linear regression, and give an illustration on the US crime data of Vandaele.

math.ST