SearcharxivSearch

arXiv subjects

Christopher M. Hans

Publications and source records attributed to Christopher M. Hans.

6 recordsLinked to original sources

Bayesian Elastic Net Regression with Structured Prior Dependence

Many regularization priors for Bayesian regression assume the regression coefficients are a priori independent. In particular this is the case for standard Bayesian treatments of the lasso and the elastic net. While independence may be reasonable in some data-analytic settings, incorporating dependence in these prior distributions provides greater modeling flexibility. This paper introduces the orthant normal distribution in its general form and shows how it can be used to structure prior dependence in the Bayesian elastic net regression model. An L1-regularized version of Zellner's g prior is introduced as a special case, creating a new link between the literature on penalized optimization and an important class of regression priors. Computation is challenging due to an intractable normalizing constant in the prior. We avoid this issue by modifying slightly a standard prior of convenience for the hyperparameters in such a way to enable simple and fast Gibbs sampling of the posterior distribution. The benefit of including structured prior dependence in the Bayesian elastic net regression model is demonstrated through simulation and a near-infrared spectroscopy data example.

stat.ME

Sampling the Bayesian Elastic Net

The Bayesian elastic net regression model is characterized by the regression coefficient prior distribution, the negative log density of which corresponds to the elastic net penalty function. While Markov chain Monte Carlo (MCMC) methods exist for sampling from the posterior of the regression coefficients given the penalty parameters, full Bayesian inference that incorporates uncertainty about the penalty parameters remains a challenge due to an intractable integrable in the posterior density function. Though sampling methods have been proposed that avoid computing this integral, all correctly-specified methods for full Bayesian inference that have appeared in the literature involve at least one "Metropolis-within-Gibbs" update, requiring tuning of proposal distributions. The computational landscape is complicated by the fact that two forms of the Bayesian elastic net prior have been introduced, and two representations (with and without data augmentation) of the prior suggest different MCMC algorithms. We review the forms and representations of the prior, discuss all combinations of these different treatments for the first time, and introduce one combination of form and representation that has yet to appear in the literature. We introduce MCMC algorithms for full Bayesian inference for all treatments of the prior. The algorithms allow for direct sampling of all parameters without any "Metropolis-within-Gibbs" steps. The key to the new approach is a careful transformation of the parameter space and an analysis of the resulting full conditional density functions that allows for efficient rejection sampling. We make empirical comparisons between our approaches and existing MCMC samplers for different data structures.

stat.CO

Empirical Bayes Model Averaging with Influential Observations: Tuning Zellner's g Prior for Predictive Robustness

The behavior of Bayesian model averaging (BMA) for the normal linear regression model in the presence of influential observations that contribute to model misfit is investigated. Remedies to attenuate the potential negative impacts of such observations on inference and prediction are proposed. The methodology is motivated by the view that well-behaved residuals and good predictive performance often go hand-in-hand. Focus is placed on regression models that use variants on Zellner's g prior. Studying the impact of various forms of model misfit on BMA predictions in simple situations points to prescriptive guidelines for "tuning" Zellner's g prior to obtain optimal predictions. The tuning of the prior distribution is obtained by considering theoretical properties that should be enjoyed by the optimal fits of the various models in the BMA ensemble. The methodology can be thought of as an "empirical Bayes" approach to modeling, as the data help to inform the specification of the prior in an attempt to attenuate the negative impact of influential cases.

stat.ME

Prediction Using a Bayesian Heteroscedastic Composite Gaussian Process

This research proposes a flexible Bayesian extension of the composite Gaussian process (CGP) model of Ba and Joseph (2012) for predicting (stationary or) non-stationary $y(\mathbf{x})$. The CGP generalizes the regression plus stationary Gaussian process (GP) model by replacing the regression term with a GP. The new model, $Y(\mathbf{x})$, can accommodate large-scale trends estimated by a global GP, local trends estimated by an independent local GP, and a third process to describe heteroscedastic data in which $Var(Y(\mathbf{x}))$ can depend on the inputs. This paper proposes a prior which ensures that the fitted global mean is smoother than the local deviations, and extends the covariance structure of the CGP to allow for differentially-weighted global and local components. A Markov chain Monte Carlo algorithm is proposed to provide posterior estimates of the parameters, including the values of the heteroscedastic variance at the training and test data locations. The posterior distribution is used to make predictions and to quantify the uncertainty of the predictions using prediction intervals. The method is illustrated using both stationary and non-stationary $y(\mathbf{x})$.

stat.ME

Block Hyper-g Priors in Bayesian Regression

The development of prior distributions for Bayesian regression has traditionally been driven by the goal of achieving sensible model selection and parameter estimation. The formalization of properties that characterize good performance has led to the development and popularization of thick tailed mixtures of g priors such as the Zellner--Siow and hyper-g priors. The properties of a particular prior are typically illuminated under limits on the likelihood or the prior. In this paper we introduce a new, conditional information asymptotic that is motivated by the common data analysis setting where at least one regression coefficient is much larger than others. We analyze existing mixtures of g priors under this limit and reveal two new behaviors, Essentially Least Squares (ELS) estimation and the Conditional Lindley's Paradox (CLP), and argue that these behaviors are, in general, undesirable. As the driver behind both of these behaviors is the use of a single, latent scale parameter that is common to all coefficients, we propose a block hyper-g prior, defined by first partitioning the covariates into groups and then placing independent hyper-g priors on the corresponding blocks of coefficients. We provide conditions under which ELS and the CLP are avoided by the new class of priors, and provide consistency results under traditional sample size asymptotics.

math.ST