SearcharxivSearch

arXiv subjects

Qian Qin

Publications and source records attributed to Qian Qin.

At least 19 recordsLinked to original sources

A global spectral gap for Metropolis-adjusted Langevin algorithm with a uniformly randomized step size

Let $\pi(\mathrm{d} x)\propto e^{-U(x)}\,\mathrm{d} x$ on $\mathbb R^d$, where $0<m\leq L<\infty$, $mI_d\preceq\nabla^2U(x)\preceq LI_d$, and $\kappa=L/m$. It is known that, under warm-start assumptions, fixed-step Metropolis-adjusted Langevin algorithm (MALA) with properly tuned step size has mixing time of order $\kappa \sqrt{d}$ up to logarithmic factors. By contrast, when the condition number is bounded away from one, no single fixed step size yields a matching spectral-gap lower bound of order $(\kappa \sqrt{d})^{-1}$ uniformly over this target class. We show that MALA with a uniformly randomized step size admits a spectral-gap lower bound of this size. At each iteration, the randomized-step MALA considered here draws $h$ uniformly from $(0,H)$ and performs one ordinary MALA transition with step size $h$. We show that, when $H$ is of order $(L\sqrt{d})^{-1}$, the right spectral gap of randomized-step MALA admits a lower bound of order \[ \frac{1}{\kappa\sqrt{d}\,[1+\log(d+1)+\log\kappa]}. \] The main new ingredient in the proof is a Cheeger-type inequality for aggregating estimates of the one-step flow of MALA out of measurable sets at various step-size scales. It allows the scale used to control the flow to depend on the set and avoids the additional loss that would result from first summing the flows and then applying the standard Cheeger inequality. This work was developed with substantial assistance from ChatGPT, which suggested the uniformly randomized-step approach, developed the principal proof arguments, and generated the simulation and Lean 4 code. The human author checked and verified the mathematical content and take full responsibility for the results.

math.ST

Spectral Gap for the Binary Fixed-Margin Swap Chain

We prove an explicit spectral-gap lower bound for the lazy swap chain on binary matrices with prescribed row and column sums. This chain is a standard sampler for fixed-margin null models in ecology, statistics, and network analysis. Kannan, Tetali, and Vempala (KTV) conjectured that it mixes rapidly for all feasible margins \citep{kannan1997simple}. We show that for every feasible set of margins on an $m\times n$ binary matrix, the lazy swap chain has spectral gap at least $$\binom{m}{2}^{-1}\binom{n}{2}^{-1}.$$ The bound is tight in the worst case. Thus, our result proves this KTV conjecture in a stronger quantitative form. The same spectral-gap bound also verifies the Mihail--Vazirani conjecture for fixed-margin 0/1-matrix polytopes. The proof gives a new route to fixed-margin sampling that avoids stability assumptions and canonical-path constructions. We compare the swap chain with a two-row heat-bath chain and use a local-to-global spectral reduction to reduce the analysis from arbitrary $m\times n$ matrices to a three-row problem. The remaining three-row inequality is then proved by separating the scalar column-count sector from the non-scalar Johnson harmonic sectors. The proof itself was generated by ChatGPT 5.5 Pro. The author's role was to pose the problem, guide the search direction, evaluate the AI-generated arguments, rewrite the proof, and take responsibility for the final form and validity of the result. The full proof of the main theorem has been formalized in Lean, and the accompanying formalization is available at the anonymous repository https://github.com/guanyangwang/ktv-swap-lean.

math.PR

Solidarity of Spectral Gaps for Component-Wise Markov Chains

Deterministic-scan and random-scan component-wise Markov chain Monte Carlo algorithms, such as Gibbs samplers and conditional Metropolis-Hastings, are popular approaches for sampling from multivariate distributions. A long-standing open question is to determine the conditions under which these algorithms have similar convergence rates. A block-wise contraction condition for the component-wise updates is used to establish a solidarity principle for the $L^2$ spectral gaps of the associated Markov chains. Specifically, under this condition, the spectral gaps of the random-scan and deterministic-scan versions of the Gibbs and component-wise chains are either simultaneously positive or simultaneously zero. Moreover, the spectral gaps differ by at most polynomial factors in the number of blocks. As an application of the general results, a deterministic-scan conditional Metropolis-adjusted Langevin algorithm (MALA) for multivariate Gaussian targets is studied. The block-wise contraction condition is combined with known spectral gap bounds for the random-scan Gibbs sampler to obtain a spectral gap bound that is polynomial in dimension. The result is used to clarify how the convergence rate of the conditional MALA depends on the precision matrix of the Gaussian target and the step sizes of the block-wise MALA updates.

math.ST

Convergence analysis of data augmentation algorithms in Bayesian lasso models with log-concave likelihoods

We study the convergence properties of a class of data augmentation algorithms targeting posterior distributions of Bayesian lasso models with log-concave likelihoods. Leveraging isoperimetric inequalities, we derive a generic convergence bound for this class of algorithms and apply it to Bayesian probit, logistic, and heteroskedastic Gaussian linear lasso models. Under feasible initializations, the mixing times for the probit and logistic models are of order $O[(p+n)^3 (pn^{1-c} + n)]$, up to logarithmic factors, where $n$ is the sample size, $p$ is the dimension of the regression coefficients, and $c \in [0,1]$ is determined by the lasso penalty parameter. The mixing time for the heteroskedastic Gaussian model is $O[n(n+p)^3 (p n^{1-c} + n)]$, up to logarithmic factors.

math.ST

Asymptotic behavior and sharp estimates for spreading fronts in a cooperative system with free boundaries

This paper investigates the dynamics of a reaction-diffusion system with two free boundaries, modeling the invasion of two cooperative species, where the free boundaries represent expanding fronts. We first analyze the long-term behavior of the system, showing that it follows a spreading-vanishing dichotomy: the two species either spread across the entire region or eventually die out. In the case of spreading, we determine the asymptotic spreading speed of the fronts by using a semi-wave system and provide sharp estimates for the moving fronts. Additionally, we show that the solution to the system converges to the corresponding semi-wave solution as time tends to infinity. These results contribute to a deeper understanding of the long-term dynamics of cooperative species in reaction-diffusion systems with free boundaries.

math.AP

Challenges in structural variant calling in low-complexity regions

Background: Structural variants (SVs) are genomic differences $\ge$50 bp in length. They remain challenging to detect even with long sequence reads, and the sources of these difficulties are not well quantified. Results: We identified 35.4 Mb of low-complexity regions (LCRs) in GRCh38. Although these regions cover only 1.2% of the genome, they contain 69.1% of confident SVs in sample HG002. Across long-read SV callers, 77.3-91.3% of erroneous SV calls occur within LCRs, with error rates increasing with LCR length. Conclusion: SVs are enriched and difficult to call in LCRs. Special care need to be taken for calling and analyzing these variants.

q-bio.GN

On spectral gap decomposition for Markov chains

Multiple works regarding convergence analysis of Markov chains have led to spectral gap decomposition formulas of the form \[ \mathrm{Gap}(S) \geq c_0 \left[\inf_z \mathrm{Gap}(Q_z)\right] \mathrm{Gap}(\bar{S}), \] where $c_0$ is a constant, $\mathrm{Gap}$ denotes the right spectral gap of a reversible Markov operator, $S$ is the Markov transition kernel (Mtk) of interest, $\bar{S}$ is an idealized or simplified version of $S$, and $\{Q_z\}$ is a collection of Mtks characterizing the differences between $S$ and $\bar{S}$. This type of relationship has been established in various contexts, including: 1. decomposition of Markov chains based on a finite cover of the state space, 2. hybrid Gibbs samplers, and 3. spectral independence and localization schemes. We show that multiple key decomposition results across these domains can be connected within a unified framework, rooted in a simple sandwich structure of $S$. Within the general framework, we establish new instances of spectral gap decomposition for hybrid hit-and-run samplers and hybrid data augmentation algorithms with two intractable conditional distributions. Additionally, we explore several other properties of the sandwich structure, and derive extensions of the spectral gap decomposition formula.

math.ST

Spectral gap bounds for reversible hybrid Gibbs chains

Hybrid Gibbs samplers represent a prominent class of approximated Gibbs algorithms that utilize Markov chains to approximate conditional distributions, with the Metropolis-within-Gibbs algorithm standing out as a well-known example. Despite their widespread use in both statistical and non-statistical applications, little is known about their convergence properties. This article introduces novel methods for establishing bounds on the convergence rates of certain reversible hybrid Gibbs samplers. In particular, we examine the convergence characteristics of hybrid random-scan Gibbs algorithms. Our analysis reveals that the absolute spectral gap of a hybrid Gibbs chain can be bounded based on the absolute spectral gap of the exact Gibbs chain and the absolute spectral gaps of the Markov chains employed for conditional distribution approximations. We also provide a convergence bound of similar flavors for hybrid data augmentation algorithms, extending existing works on the topic. The general bounds are applied to three examples: a random-scan Metropolis-within-Gibbs sampler, random-scan Gibbs samplers with block updates, and a hybrid slice sampler.

math.ST

Geometric ergodicity of trans-dimensional Markov chain Monte Carlo algorithms

This article studies the convergence properties of trans-dimensional MCMC algorithms when the total number of models is finite. It is shown that, for reversible and some non-reversible trans-dimensional Markov chains, under mild conditions, geometric convergence is guaranteed if the Markov chains associated with the within-model moves are geometrically ergodic. This result is proved in an $L^2$ framework using the technique of Markov chain decomposition. While the technique was previously developed for reversible chains, this work extends it to the point that it can be applied to some commonly used non-reversible chains. The theory herein is applied to reversible jump algorithms for three Bayesian models: a probit regression with variable selection, a Gaussian mixture model with unknown number of components, and an autoregression with Laplace errors and unknown model order.

math.ST

A phase transition in sampling from Restricted Boltzmann Machines

Restricted Boltzmann Machines are a class of undirected graphical models that play a key role in deep learning and unsupervised learning. In this study, we prove a phase transition phenomenon in the mixing time of the Gibbs sampler for a one-parameter Restricted Boltzmann Machine. Specifically, the mixing time varies logarithmically, polynomially, and exponentially with the number of vertices depending on whether the parameter $c$ is above, equal to, or below a critical value $c_\star\approx-5.87$. A key insight from our analysis is the link between the Gibbs sampler and a dynamical system, which we utilize to quantify the former based on the behavior of the latter. To study the critical case $c= c_\star$, we develop a new isoperimetric inequality for the sampler's stationary distribution by showing that the distribution is nearly log-concave.

cs.LG

Convergence Bounds for Monte Carlo Markov Chains

This review paper, written for the second edition of the Handbook of Markov Chain Monte Carlo, provides an introduction to the study of convergence analysis for Markov chain Monte Carlo (MCMC), aimed at researchers new to the field. We focus on methods for constructing bounds on the distance between the distribution of a Markov chain at a given time and its stationary distribution. Two widely-used approaches are explored: the coupling method and the L2 theory of Markov chains. For the latter, we emphasize techniques based on conductance and isoperimetric inequalities. Additionally, we briefly discuss strategies for identifying slow convergence in Markov chains.

math.ST

Neural-g: A Deep Learning Framework for Mixing Density Estimation

Mixing (or prior) density estimation is an important problem in machine learning and statistics, especially in empirical Bayes $g$-modeling where accurately estimating the prior is necessary for making good posterior inferences. In this paper, we propose neural-$g$, a new neural network-based estimator for $g$-modeling. Neural-$g$ uses a softmax output layer to ensure that the estimated prior is a valid probability density. Under default hyperparameters, we show that neural-$g$ is very flexible and capable of capturing many unknown densities, including those with flat regions, heavy tails, and/or discontinuities. In contrast, existing methods struggle to capture all of these prior shapes. We provide justification for neural-$g$ by establishing a new universal approximation theorem regarding the capability of neural networks to learn arbitrary probability mass functions. To accelerate convergence of our numerical implementation, we utilize a weighted average gradient descent approach to update the network parameters. Finally, we extend neural-$g$ to multivariate prior density estimation. We illustrate the efficacy of our approach through simulations and analyses of real datasets. A software package to implement neural-$g$ is publicly available at https://github.com/shijiew97/neuralG.

stat.ML

Multivariate strong invariance principle and uncertainty assessment for time in-homogeneous cyclic MCMC samplers

Time in-homogeneous cyclic Markov chain Monte Carlo (MCMC) samplers, including deterministic scan Gibbs samplers and Metropolis within Gibbs samplers, are extensively used for sampling from multi-dimensional distributions. We establish a multivariate strong invariance principle (SIP) for Markov chains associated with these samplers. The rate of this SIP essentially aligns with the tightest rate available for time homogeneous Markov chains. The SIP implies the strong law of large numbers (SLLN) and the central limit theorem (CLT), and plays an essential role in uncertainty assessments. Using the SIP, we give conditions under which the multivariate batch means estimator for estimating the covariance matrix in the multivariate CLT is strongly consistent. Additionally, we provide conditions for a multivariate fixed volume sequential termination rule, which is associated with the concept of effective sample size (ESS), to be asymptotically valid. Our uncertainty assessment tools are demonstrated through various numerical experiments.

stat.CO

Analysis of two-component Gibbs samplers using the theory of two projections

The theory of two projections is utilized to study two-component Gibbs samplers. Through this theory, previously intractable problems regarding the asymptotic variances of two-component Gibbs samplers are reduced to elementary matrix algebra exercises. It is found that in terms of asymptotic variance, the two-component random-scan Gibbs sampler is never much worse, and could be considerably better than its deterministic-scan counterpart, provided that the selection probability is appropriately chosen. This is especially the case when there is a large discrepancy in computation cost between the two components. The result contrasts with the known fact that the deterministic-scan version has a faster convergence rate, which can also be derived from the method herein. On the other hand, a modified version of the deterministic-scan sampler that accounts for computation cost can outperform the random-scan version.

math.ST

Convergence Analysis of Data Augmentation Algorithms for Bayesian Robust Multivariate Linear Regression with Incomplete Data

Gaussian mixtures are commonly used for modeling heavy-tailed error distributions in robust linear regression. Combining the likelihood of a multivariate robust linear regression model with a standard improper prior distribution yields an analytically intractable posterior distribution that can be sampled using a data augmentation algorithm. When the response matrix has missing entries, there are unique challenges to the application and analysis of the convergence properties of the algorithm. Conditions for geometric ergodicity are provided when the incomplete data have a "monotone" structure. In the absence of a monotone structure, an intermediate imputation step is necessary for implementing the algorithm. In this case, we provide sufficient conditions for the algorithm to be Harris ergodic. Finally, we show that, when there is a monotone structure and intermediate imputation is unnecessary, intermediate imputation slows the convergence of the underlying Monte Carlo Markov chain, while post hoc imputation does not. An R package for the data augmentation algorithm is provided.

math.ST

Spectral Telescope: Convergence Rate Bounds for Random-Scan Gibbs Samplers Based on a Hierarchical Structure

Random-scan Gibbs samplers possess a natural hierarchical structure. The structure connects Gibbs samplers targeting higher dimensional distributions to those targeting lower dimensional ones. This leads to a quasi-telescoping property of their spectral gaps. Based on this property, we derive three new bounds on the spectral gaps and convergence rates of Gibbs samplers on general domains. The three bounds relate a chain's spectral gap to, respectively, the correlation structure of the target distribution, a class of random walk chains, and a collection of influence matrices. Notably, one of our results generalizes the technique of spectral independence, which has received considerable attention for its success on finite domains, to general state spaces. We illustrate our methods through a sampler targeting the uniform distribution on a corner of an $n$-cube.

math.PR

Convergence Rates of Two-Component MCMC Samplers

Component-wise MCMC algorithms, including Gibbs and conditional Metropolis-Hastings samplers, are commonly used for sampling from multivariate probability distributions. A long-standing question regarding Gibbs algorithms is whether a deterministic-scan (systematic-scan) sampler converges faster than its random-scan counterpart. We answer this question when the samplers involve two components by establishing an exact quantitative relationship between the $L^2$ convergence rates of the two samplers. The relationship shows that the deterministic-scan sampler converges faster. We also establish qualitative relations among the convergence rates of two-component Gibbs samplers and some conditional Metropolis-Hastings variants. For instance, it is shown that if some two-component conditional Metropolis-Hastings samplers are geometrically ergodic, then so are the associated Gibbs samplers.

math.ST

Geometric convergence bounds for Markov chains in Wasserstein distance based on generalized drift and contraction conditions

Let $(X_n)_{n=0}^\infty$ denote a Markov chain on a Polish space that has a stationary distribution $\varpi$. This article concerns upper bounds on the Wasserstein distance between the distribution of $X_n$ and $\varpi$. In particular, an explicit geometric bound on the distance to stationarity is derived using generalized drift and contraction conditions whose parameters vary across the state space. These new types of drift and contraction allow for sharper convergence bounds than the standard versions, whose parameters are constant. Application of the result is illustrated in the context of a non-linear autoregressive process and a Gibbs algorithm for a random effects model.

math.PR