Searcharxiv⌕ Search

arXiv subjects

Yves F. Atchadé

Publications and source records attributed to Yves F. Atchadé.

11 recordsLinked to original sources

Unbiased Markov Chain Monte Carlo: what, why, and how

This document presents methods to remove the initialization or burn-in bias from Markov chain Monte Carlo (MCMC) estimates, with consequences on parallel computing, convergence diagnostics and performance assessment. The document is written as an introduction to these methods for MCMC users. Some theoretical results are mentioned, but the focus is on the methodology.

stat.ME↗

Approximate spectral gaps for Markov chains mixing times in high dimensions

This paper introduces a concept of approximate spectral gap to analyze the mixing time of Markov Chain Monte Carlo (MCMC) algorithms for which the usual spectral gap is degenerate or almost degenerate. We use the idea to analyze a class of MCMC algorithms to sample from mixtures of densities. As an application we study the mixing time of a Gibbs sampler for variable selection in linear regression models. Under some regularity conditions on the signal and the design matrix of the regression problem, we show that for well-chosen initial distributions the mixing time of the Gibbs sampler is polynomial in the dimension of the space.

stat.CO↗

Markov chain Monte Carlo algorithms with sequential proposals

We explore a general framework in Markov chain Monte Carlo (MCMC) sampling where sequential proposals are tried as a candidate for the next state of the Markov chain. This sequential-proposal framework can be applied to various existing MCMC methods, including Metropolis-Hastings algorithms using random proposals and methods that use deterministic proposals such as Hamiltonian Monte Carlo (HMC) or the bouncy particle sampler. Sequential-proposal MCMC methods construct the same Markov chains as those constructed by the delayed rejection method under certain circumstances. In the context of HMC, the sequential-proposal approach has been proposed as extra chance generalized hybrid Monte Carlo (XCGHMC). We develop two novel methods in which the trajectories leading to proposals in HMC are automatically tuned to avoid doubling back, as in the No-U-Turn sampler (NUTS). The numerical efficiency of these new methods compare favorably to the NUTS. We additionally show that the sequential-proposal bouncy particle sampler enables the constructed Markov chain to pass through regions of low target density and thus facilitates better mixing of the chain when the target density is multimodal.

stat.CO↗

Unbiased Markov chain Monte Carlo with couplings

Markov chain Monte Carlo (MCMC) methods provide consistent of integrals as the number of iterations goes to infinity. MCMC estimators are generally biased after any fixed number of iterations. We propose to remove this bias by using couplings of Markov chains together with a telescopic sum argument of Glynn and Rhee (2014). The resulting unbiased estimators can be computed independently in parallel. We discuss practical couplings for popular MCMC algorithms. We establish the theoretical validity of the proposed estimators and study their efficiency relative to the underlying MCMC algorithms. Finally, we illustrate the performance and limitations of the method on toy examples, on an Ising model around its critical temperature, on a high-dimensional variable selection problem, and on an approximation of the cut distribution arising in Bayesian inference for models made of multiple modules.

stat.ME↗

On the contraction properties of some high-dimensional quasi-posterior distributions

We study the contraction properties of a quasi-posterior distribution $\checkΠ_{n,d}$ obtained by combining a quasi-likelihood function and a sparsity inducing prior distribution on $\rset^d$, as both $n$ (the sample size), and $d$ (the dimension of the parameter) increase. We derive some general results that highlight a set of sufficient conditions under which $\checkΠ_{n,d}$ puts increasingly high probability on sparse subsets of $\rset^d$, and contracts towards the true value of the parameter. We apply these results to the analysis of logistic regression models, and binary graphical models, in high-dimensional settings. For the logistic regression model, we shows that for well-behaved design matrices, the posterior distribution contracts at the rate $O(\sqrt{s_\star\log(d)/n})$, where $s_\star$ is the number of non-zero components of the parameter. For the binary graphical model, under some regularity conditions, we show that a quasi-posterior analog of the neighborhood selection of \cite{meinshausen06} contracts in the Frobenius norm at the rate $O(\sqrt{(p+S)\log(p)/n})$, where $p$ is the number of nodes, and $S$ the number of edges of the true graph.

math.ST↗

A Moreau-Yosida approximation scheme for a class of high-dimensional posterior distributions

Exact-sparsity inducing prior distributions in Bayesian analysis typically lead to posterior distributions that are very challenging to handle by standard Markov Chain Monte Carlo (MCMC) methods, particular in high-dimensional models with large number of parameters. We propose a methodology to derive smooth approximations of such posterior distributions that are, in some cases, easier to handle by standard MCMC methods. The approximation is obtained from the forward-backward approximation of the Moreau-Yosida regularization of the negative log-density. We show that the derived approximation is within $O(\sqrtγ)$ of the true posterior distribution in the $β$-metric, where $γ>0$ is a user-controlled parameter that defines the approximation. We illustrate the method with a high-dimensional linear regression model.

math.ST↗

Markov Chain Monte Carlo confidence intervals

For a reversible and ergodic Markov chain $\{X_n,n\geq0\}$ with invariant distribution $π$, we show that a valid confidence interval for $π(h)$ can be constructed whenever the asymptotic variance $σ^2_P(h)$ is finite and positive. We do not impose any additional condition on the convergence rate of the Markov chain. The confidence interval is derived using the so-called fixed-b lag-window estimator of $σ_P^2(h)$. We also derive a result that suggests that the proposed confidence interval procedure converges faster than classical confidence interval procedures based on the Gaussian distribution and standard central limit theorems for Markov chains.

math.ST↗

Scalable Computation of Regularized Precision Matrices via Stochastic Optimization

We consider the problem of computing a positive definite $p \times p$ inverse covariance matrix aka precision matrix $θ=(θ_{ij})$ which optimizes a regularized Gaussian maximum likelihood problem, with the elastic-net regularizer $\sum_{i,j=1}^{p} λ(α|θ_{ij}| + \frac{1}{2}(1- α) θ_{ij}^2),$ with regularization parameters $α\in [0,1]$ and $λ>0$. The associated convex semidefinite optimization problem is notoriously difficult to scale to large problems and has demanded significant attention over the past several years. We propose a new algorithmic framework based on stochastic proximal optimization (on the primal problem) that can be used to obtain near optimal solutions with substantial computational savings over deterministic algorithms. A key challenge of our work stems from the fact that the optimization problem being investigated does not satisfy the usual assumptions required by stochastic gradient methods. Our proposal has (a) computational guarantees and (b) scales well to large problems, even if the solution is not too sparse; thereby, enhancing the scope of regularized maximum likelihood problems to many large-scale problems of contemporary interest. An important aspect of our proposal is to bypass the \emph{deterministic} computation of a matrix inverse by drawing random samples from a suitable multivariate Gaussian distribution.

math.ST↗

Kernel estimators of asymptotic variance for adaptive Markov chain Monte Carlo

We study the asymptotic behavior of kernel estimators of asymptotic variances (or long-run variances) for a class of adaptive Markov chains. The convergence is studied both in $L^p$ and almost surely. The results also apply to Markov chains and improve on the existing literature by imposing weaker conditions. We illustrate the results with applications to the $\operatorname {GARCH}(1,1)$ Markov model and to an adaptive MCMC algorithm for Bayesian logistic regression.

math.PR↗

A cautionary tale on the efficiency of some adaptive Monte Carlo schemes

There is a growing interest in the literature for adaptive Markov chain Monte Carlo methods based on sequences of random transition kernels $\{P_n\}$ where the kernel $P_n$ is allowed to have an invariant distribution $π_n$ not necessarily equal to the distribution of interest $π$ (target distribution). These algorithms are designed such that as $n\to\infty$, $P_n$ converges to $P$, a kernel that has the correct invariant distribution $π$. Typically, $P$ is a kernel with good convergence properties, but one that cannot be directly implemented. It is then expected that the algorithm will inherit the good convergence properties of $P$. The equi-energy sampler of [Ann. Statist. 34 (2006) 1581--1619] is an example of this type of adaptive MCMC. We show in this paper that the asymptotic variance of this type of adaptive MCMC is always at least as large as the asymptotic variance of the Markov chain with transition kernel $P$. We also show by simulation that the difference can be substantial.

stat.CO↗

Discussion of "Equi-energy sampler" by Kou, Zhou and Wong

We congratulate Samuel Kou, Qing Zhou and Wing Wong [math.ST/0507080] (referred to subsequently as KZW) for this beautifully written paper, which opens a new direction in Monte Carlo computation. This discussion has two parts. First, we describe a very closely related method, multicanonical sampling (MCS), and report a simulation example that compares the equi-energy (EE) sampler with MCS. Overall, we found the two algorithms to be of comparable efficiency for the simulation problem considered. In the second part, we develop some additional convergence results for the EE sampler.

math.ST↗