SearcharxivSearch

arXiv subjects

James P. Hobert

Publications and source records attributed to James P. Hobert.

At least 19 recordsLinked to original sources

Extensions of the solidarity principle of the spectral gap for Gibbs samplers to their blocked and collapsed variants

Connections of a spectral nature are formed between Gibbs samplers and their blocked and collapsed variants. The solidarity principle of the spectral gap for full Gibbs samplers is generalized to different cycles and mixtures of Gibbs steps. This generalized solidarity principle is employed to establish that every cycle and mixture of Gibbs steps, which includes blocked Gibbs samplers and collapsed Gibbs samplers, inherits a spectral gap from a full Gibbs sampler. Exact relations between the spectra corresponding to blocked and collapsed variants of a Gibbs sampler are also established. An example is given to show that a blocked or collapsed Gibbs sampler does not in general inherit geometric ergodicity or a spectral gap from another blocked or collapsed Gibbs sampler.

stat.CO

The data augmentation algorithm

The data augmentation (DA) algorithms are popular Markov chain Monte Carlo (MCMC) algorithms often used for sampling from intractable probability distributions. This review article comprehensively surveys DA MCMC algorithms, highlighting their theoretical foundations, methodological implementations, and diverse applications in frequentist and Bayesian statistics. The article discusses tools for studying the convergence properties of DA algorithms. Furthermore, it contains various strategies for accelerating the speed of convergence of the DA algorithms, different extensions of DA algorithms and outlines promising directions for future research. This paper aims to serve as a resource for researchers and practitioners seeking to leverage data augmentation techniques in MCMC algorithms by providing key insights and synthesizing recent developments.

stat.CO

On the convergence rate of the "out-of-order" block Gibbs sampler

It is shown that a seemingly harmless reordering of the steps in a block Gibbs sampler can actually invalidate the algorithm. In particular, the Markov chain that is simulated by the "out-of-order" block Gibbs sampler does not have the correct invariant probability distribution. However, despite having the wrong invariant distribution, the Markov chain converges at the same rate as the original block Gibbs Markov chain. More specifically, it is shown that either both Markov chains are geometrically ergodic (with the same geometric rate of convergence), or neither one is. These results are important from a practical standpoint because the (invalid) out-of-order algorithm may be easier to analyze than the (valid) block Gibbs sampler (see, e.g., Yang and Rosenthal [2019]).

math.ST

Dimension free convergence rates for Gibbs samplers for Bayesian linear mixed models

The emergence of big data has led to a growing interest in so-called convergence complexity analysis, which is the study of how the convergence rate of a Monte Carlo Markov chain (for an intractable Bayesian posterior distribution) scales as the underlying data set grows in size. Convergence complexity analysis of practical Monte Carlo Markov chains on continuous state spaces is quite challenging, and there have been very few successful analyses of such chains. One fruitful analysis was recently presented by Qin and Hobert (2021b), who studied a Gibbs sampler for a simple Bayesian random effects model. These authors showed that, under regularity conditions, the geometric convergence rate of this Gibbs sampler converges to zero as the data set grows in size. It is shown herein that similar behavior is exhibited by Gibbs samplers for more general Bayesian models that possess both random effects and traditional continuous covariates, the so-called mixed models. The analysis employs the Wasserstein-based techniques introduced by Qin and Hobert (2021b).

math.ST

Approximating the Spectral Gap of the Pólya-Gamma Gibbs Sampler

The self-adjoint, positive Markov operator defined by the Pólya-Gamma Gibbs sampler (under a proper normal prior) is shown to be trace-class, which implies that all non-zero elements of its spectrum are eigenvalues. Consequently, the spectral gap is $1-λ_*$, where $λ_* \in [0,1)$ is the second largest eigenvalue. A method of constructing an asymptotically valid confidence interval for an upper bound on $λ_*$ is developed by adapting the classical Monte Carlo technique of Qin et al. (2019) to the Pólya-Gamma Gibbs sampler. The results are illustrated using the German credit data. It is also shown that, in general, uniform ergodicity does not imply the trace-class property, nor does the trace-class property imply uniform ergodicity.

math.ST

Geometric convergence bounds for Markov chains in Wasserstein distance based on generalized drift and contraction conditions

Let $(X_n)_{n=0}^\infty$ denote a Markov chain on a Polish space that has a stationary distribution $\varpi$. This article concerns upper bounds on the Wasserstein distance between the distribution of $X_n$ and $\varpi$. In particular, an explicit geometric bound on the distance to stationarity is derived using generalized drift and contraction conditions whose parameters vary across the state space. These new types of drift and contraction allow for sharper convergence bounds than the standard versions, whose parameters are constant. Application of the result is illustrated in the context of a non-linear autoregressive process and a Gibbs algorithm for a random effects model.

math.PR

Wasserstein-based methods for convergence complexity analysis of MCMC with applications

Over the last 25 years, techniques based on drift and minorization (d&m) have been mainstays in the convergence analysis of MCMC algorithms. However, results presented herein suggest that d&m may be less useful in the emerging area of convergence complexity analysis, which is the study of how the convergence behavior of Monte Carlo Markov chains scale with sample size, $n$, and/or number of covariates, $p$. The problem appears to be that minorization can become a serious liability as dimension increases. Alternative methods for constructing convergence rate bounds (with respect to total variation distance) that do not require minorization are investigated. Based on Wasserstein distances and random mappings, these methods can produce bounds that are substantially more robust to increasing dimension than those based on d&m. The Wasserstein-based bounds are used to develop strong convergence complexity results for MCMC algorithms used in Bayesian probit regression and random effects models in the challenging asymptotic regime where $n$ and $p$ are both large.

math.ST

On the limitations of single-step drift and minorization in Markov chain convergence analysis

Over the last three decades, there has been a considerable effort within the applied probability community to develop techniques for bounding the convergence rates of general state space Markov chains. Most of these results assume the existence of drift and minorization (d\&m) conditions. It has often been observed that convergence rate bounds based on single-step d\&m tend to be overly conservative, especially in high-dimensional situations. This article builds a framework for studying this phenomenon. It is shown that any convergence rate bound based on a set of d\&m conditions cannot do better than a certain unknown optimal bound. Strategies are designed to put bounds on the optimal bound itself, and this allows one to quantify the extent to which a d\&m-based convergence rate bound can be sharp. The new theory is applied to several examples, including a Gaussian autoregressive process (whose true convergence rate is known), and a Metropolis adjusted Langevin algorithm. The results strongly suggest that convergence rate bounds based on single-step d\&m conditions are quite inadequate in high-dimensional settings.

math.PR

On the convergence complexity of Gibbs samplers for a family of simple Bayesian random effects models

The emergence of big data has led to so-called convergence complexity analysis, which is the study of how Markov chain Monte Carlo (MCMC) algorithms behave as the sample size, $n$, and/or the number of parameters, $p$, in the underlying data set increase. This type of analysis is often quite challenging, in part because existing results for fixed $n$ and $p$ are simply not sharp enough to yield good asymptotic results. One of the first convergence complexity results for an MCMC algorithm on a continuous state space is due to Yang and Rosenthal (2019), who established a mixing time result for a Gibbs sampler (for a simple Bayesian random effects model) that was introduced and studied by Rosenthal (1996). The asymptotic behavior of the spectral gap of this Gibbs sampler is, however, still unknown. We use a recently developed simulation technique (Qin et. al., 2019) to provide substantial numerical evidence that the gap is bounded away from 0 as $n \rightarrow \infty$. We also establish a pair of rigorous convergence complexity results for two different Gibbs samplers associated with a generalization of the random effects model considered by Rosenthal (1996). Our results show that, under strong regularity conditions, the spectral gaps of these Gibbs samplers converge to 1 as the sample size increases.

math.ST

A Hybrid Scan Gibbs Sampler for Bayesian Models with Latent Variables

Gibbs sampling is a widely popular Markov chain Monte Carlo algorithm that can be used to analyze intractable posterior distributions associated with Bayesian hierarchical models. There are two standard versions of the Gibbs sampler: The systematic scan (SS) version, where all variables are updated at each iteration, and the random scan (RS) version, where a single, randomly selected variable is updated at each iteration. The literature comparing the theoretical properties of SS and RS Gibbs samplers is reviewed, and an alternative hybrid scan Gibbs sampler is introduced, which is particularly well suited to Bayesian models with latent variables. The word "hybrid" reflects the fact that the scan used within this algorithm has both systematic and random elements. Indeed, at each iteration, one updates the entire set of latent variables, along with a randomly chosen block of the remaining variables. The hybrid scan (HS) Gibbs sampler has important advantages over the two standard scan Gibbs samplers. Firstly, the HS algorithm is often easier to analyze from a theoretical standpoint. In particular, it can be much easier to establish the geometric ergodicity of a HS Gibbs Markov chain than to do the same for the corresponding SS and RS versions. Secondly, the sandwich methodology developed in Hobert and Marchev (2008), which is also reviewed, can be applied to the HS Gibbs algorithm (but not to the standard scan Gibbs samplers). It is shown that, under weak regularity conditions, adding sandwich steps to the HS Gibbs sampler always results in a theoretically superior algorithm. Three specific Bayesian hierarchical models of varying complexity are used to illustrate the results. One is a simple location-scale model for data from the Student's $t$ distribution, which is used as a pedagogical tool. The other two are sophisticated, yet practical Bayesian regression models.

math.ST

Estimating the spectral gap of a trace-class Markov operator

The utility of a Markov chain Monte Carlo algorithm is, in large part, determined by the size of the spectral gap of the corresponding Markov operator. However, calculating (and even approximating) the spectral gaps of practical Monte Carlo Markov chains in statistics has proven to be an extremely difficult and often insurmountable task, especially when these chains move on continuous state spaces. In this paper, a method for accurate estimation of the spectral gap is developed for general state space Markov chains whose operators are non-negative and trace-class. The method is based on the fact that the second largest eigenvalue (and hence the spectral gap) of such operators can be bounded above and below by simple functions of the power sums of the eigenvalues. These power sums often have nice integral representations. A classical Monte Carlo method is proposed to estimate these integrals, and a simple sufficient condition for finite variance is provided. This leads to asymptotically valid confidence intervals for the second largest eigenvalue (and the spectral gap) of the Markov operator. In contrast with previously existing techniques, our method is not based on a near-stationary version of the Markov chain, which, paradoxically, cannot be obtained in a principled manner without bounds on the spectral gap. On the other hand, it can be quite expensive from a computational standpoint. The efficiency of the method is studied both theoretically and empirically.

math.ST

Convergence complexity analysis of Albert and Chib's algorithm for Bayesian probit regression

The use of MCMC algorithms in high dimensional Bayesian problems has become routine. This has spurred so-called convergence complexity analysis, the goal of which is to ascertain how the convergence rate of a Monte Carlo Markov chain scales with sample size, $n$, and/or number of covariates, $p$. This article provides a thorough convergence complexity analysis of Albert and Chib's (1993) data augmentation algorithm for the Bayesian probit regression model. The main tools used in this analysis are drift and minorization conditions. The usual pitfalls associated with this type of analysis are avoided by utilizing centered drift functions, which are minimized in high posterior probability regions, and by using a new technique to suppress high-dimensionality in the construction of minorization conditions. The main result is that the geometric convergence rate of the underlying Markov chain is bounded below 1 both as $n \rightarrow \infty$ (with $p$ fixed), and as $p \rightarrow \infty$ (with $n$ fixed). Furthermore, the first computable bounds on the total variation distance to stationarity are byproducts of the asymptotic analysis.

math.ST

Fast Monte Carlo Markov chains for Bayesian shrinkage models with random effects

When performing Bayesian data analysis using a general linear mixed model, the resulting posterior density is almost always analytically intractable. However, if proper conditionally conjugate priors are used, there is a simple two-block Gibbs sampler that is geometrically ergodic in nearly all practical settings, including situations where $p > n$ (Abrahamsen and Hobert, 2017). Unfortunately, the (conditionally conjugate) multivariate normal prior on $β$ does not perform well in the high-dimensional setting where $p \gg n$. In this paper, we consider an alternative model in which the multivariate normal prior is replaced by the normal-gamma shrinkage prior developed by Griffin and Brown (2010). This change leads to a much more complex posterior density, and we develop a simple MCMC algorithm for exploring it. This algorithm, which has both deterministic and random scan components, is easier to analyze than the more obvious three-step Gibbs sampler. Indeed, we prove that the new algorithm is geometrically ergodic in most practical settings.

math.ST

Convergence analysis of block Gibbs samplers for Bayesian linear mixed models with $p>N$

Exploration of the intractable posterior distributions associated with Bayesian versions of the general linear mixed model is often performed using Markov chain Monte Carlo. In particular, if a conditionally conjugate prior is used, then there is a simple two-block Gibbs sampler available. Román and Hobert [Linear Algebra Appl. 473 (2015) 54-77] showed that, when the priors are proper and the $X$ matrix has full column rank, the Markov chains underlying these Gibbs samplers are nearly always geometrically ergodic. In this paper, Román and Hobert's (2015) result is extended by allowing improper priors on the variance components, and, more importantly, by removing all assumptions on the $X$ matrix. So, not only is $X$ allowed to be (column) rank deficient, which provides additional flexibility in parameterizing the fixed effects, it is also allowed to have more columns than rows, which is necessary in the increasingly important situation where $p>N$. The full rank assumption on $X$ is at the heart of Román and Hobert's (2015) proof. Consequently, the extension to unrestricted $X$ requires a substantially different analysis.

math.ST

Trace-class Monte Carlo Markov Chains for Bayesian Multivariate Linear Regression with Non-Gaussian Errors

Let $π$ denote the intractable posterior density that results when the likelihood from a multivariate linear regression model with errors from a scale mixture of normals is combined with the standard non-informative prior. There is a simple data augmentation algorithm (based on latent data from the mixing density) that can be used to explore $π$. Let $h(\cdot)$ and $d$ denote the mixing density and the dimension of the regression model, respectively. Hobert et al. (2016) [arXiv:1506.03113v2] have recently shown that, if $h$ converges to 0 at the origin at an appropriate rate, and $\int_0^\infty u^{\frac{d}{2}} \, h(u) \, du < \infty$, then the Markov chains underlying the DA algorithm and an alternative Haar PX-DA algorithm are both geometrically ergodic. In fact, something much stronger than geometric ergodicity often holds. Indeed, it is shown in this paper that, under simple conditions on $h$, the Markov operators defined by the DA and Haar PX-DA Markov chains are trace-class, i.e., compact with summable eigenvalues. Many of the mixing densities that satisfy Hobert et al.'s (2016) conditions also satisfy the new conditions developed in this paper. Thus, for this set of mixing densities, the new results provide a substantial strengthening of Hobert et al.'s (2016) conclusion without any additional assumptions. For example, Hobert et al. (2016) showed that the DA and Haar PX-DA Markov chains are geometrically ergodic whenever the mixing density is generalized inverse Gaussian, log-normal, Fréchet (with shape parameter larger than $d/2$), or inverted gamma (with shape parameter larger than $d/2$). The results in this paper show that, in each of these cases, the DA and Haar PX-DA Markov operators are, in fact, trace-class.

math.ST

Convergence Analysis of the Data Augmentation Algorithm for Bayesian Linear Regression with Non-Gaussian Errors

Gaussian errors are sometimes inappropriate in a multivariate linear regression setting because, for example, the data contain outliers. In such situations, it is often assumed that the error density is a scale mixture of multivariate normal densities that takes the form $f(\varepsilon) = \int_0^\infty |Σ|^{-\frac{1}{2}} u^{\frac{d}{2}} \, ϕ_d \big( Σ^{-\frac{1}{2}} \sqrt{u} \, \varepsilon \big) \, h(u) \, du$, where $d$ is the dimension of the response, $ϕ_d(\cdot)$ is the standard $d$-variate normal density, $Σ$ is an unknown $d \times d$ positive definite scale matrix, and $h(\cdot)$ is some fixed mixing density. Combining this alternative regression model with a default prior on the unknown parameters results in a highly intractable posterior density. Fortunately, there is a simple data augmentation (DA) algorithm and a corresponding Haar PX-DA algorithm that can be used to explore this posterior. This paper provides conditions (on $h$) for geometric ergodicity of the Markov chains underlying these Markov chain Monte Carlo (MCMC) algorithms. These results are extremely important from a practical standpoint because geometric ergodicity guarantees the existence of the central limit theorems that form the basis of all the standard methods of calculating valid asymptotic standard errors for MCMC-based estimators. The main result is that, if $h$ converges to 0 at the origin at an appropriate rate, and $\int_0^\infty u^{\frac{d}{2}} \, h(u) \, du < \infty$, then the DA and Haar PX-DA Markov chains are both geometrically ergodic. This result is quite far-reaching. For example, it implies the geometric ergodicity of the DA and Haar PX-DA Markov chains whenever $h$ is generalized inverse Gaussian, log-normal, inverted gamma (with shape parameter larger than $d/2$), or Fréchet (with shape parameter larger than $d/2$).

math.ST

On the Data Augmentation Algorithm for Bayesian Multivariate Linear Regression with Non-Gaussian Errors

Let $π$ denote the intractable posterior density that results when the likelihood from a multivariate linear regression model with errors from a scale mixture of normals is combined with the standard non-informative prior. There is a simple data augmentation algorithm (based on latent data from the mixing density) that can be used to explore $π$. Hobert et al. (2015) [arXiv:1506.03113v1] recently performed a convergence rate analysis of the Markov chain underlying this MCMC algorithm in the special case where the regression model is univariate. These authors provide simple sufficient conditions (on the mixing density) for geometric ergodicity of the Markov chain. In this note, we extend Hobert et al.'s (2015) result to the multivariate case.

math.ST

Convergence analysis of the Gibbs sampler for Bayesian general linear mixed models with improper priors

Bayesian analysis of data from the general linear mixed model is challenging because any nontrivial prior leads to an intractable posterior density. However, if a conditionally conjugate prior density is adopted, then there is a simple Gibbs sampler that can be employed to explore the posterior density. A popular default among the conditionally conjugate priors is an improper prior that takes a product form with a flat prior on the regression parameter, and so-called power priors on each of the variance components. In this paper, a convergence rate analysis of the corresponding Gibbs sampler is undertaken. The main result is a simple, easily-checked sufficient condition for geometric ergodicity of the Gibbs-Markov chain. This result is close to the best possible result in the sense that the sufficient condition is only slightly stronger than what is required to ensure posterior propriety. The theory developed in this paper is extremely important from a practical standpoint because it guarantees the existence of central limit theorems that allow for the computation of valid asymptotic standard errors for the estimates computed using the Gibbs sampler.

math.ST