SearcharxivSearch

arXiv subjects

Xiyun Jiao

Publications and source records attributed to Xiyun Jiao.

9 recordsLinked to original sources

HyperMC: Multi-Fidelity Hyperparameter Tuning for Stochastic Gradient MCMC

Stochastic gradient Markov chain Monte Carlo (SGMCMC) methods enable scalable Bayesian inference, but their performance depends strongly on hyperparameters such as the step size, mini-batch size, and number of leapfrog steps. Since most SGMCMC algorithms lack a Metropolis-Hastings acceptance rate, standard acceptance-based tuning methods are not directly applicable. We propose HyperMC, a multi-fidelity tuning framework that combines Hyperband-style resource allocation with kernel Stein discrepancy (KSD) evaluation. By running multiple successive-halving brackets, HyperMC balances broad exploration of a continuous hyperparameter space with increasingly accurate evaluation of promising configurations under a fixed computational budget. We further introduce Robust HyperMC, which uses global grid initialization followed by elite-guided local refinement to reduce sensitivity to random candidate generation and noisy finite-budget evaluations. Under suitable approximation and concentration conditions for the estimated KSD, we establish that the successive-halving component selects a near-optimal configuration among the sampled candidates with high probability and derive a sufficient computational budget for successful selection. Experiments on logistic regression, probabilistic matrix factorization, and Bayesian neural networks show that HyperMC improves posterior approximation or predictive calibration relative to MAMBA, grid search, and heuristic baselines, while Robust HyperMC yields more stable and reproducible tuning results.

stat.ML

Structured Dimension-Matched Joint Variational Transdimensional Inference

Bayesian model selection couples a discrete model indicator with a model-specific continuous parameter space. We introduce structured dimension-matched variational transdimensional inference (SM-VTI) for finite enumerable model spaces. A rooted construction graph expresses a model as a sequence of local stop/child decisions. Each typed edge compiles a declared scientific parent-child edit into an exact native-coordinate dimension-matching lifting; an edge-conditioned flow then learns the residual continuous transport. The resulting local policy and conditional flow define one direct joint variational distribution, without embedding every model in a saturated maximum-dimensional surrogate. We derive its exact path density and optimize the joint reverse-KL objective. On a controlled 15-model target, SM-VTI-Joint recovers terminal masses, local actions, and nonlinear conditional geometry. On a 128-model misspecified robust variable-selection problem, a 10-data-set nearly parameter-matched affine comparison with AVTI shows stronger early model-mass recovery and competitive final joint accuracy under the same target-evaluation budget.

stat.CO

Mixing efficiency of trans-model Markov chain Monte Carlo algorithms with applications in Bayesian phylogenetics

Trans-model Markov chain Monte Carlo (MCMC) algorithms are widely used in Bayesian inference, and are particularly important in Bayesian phylogenetics where phylogenetic trees represent different statistical models. While the algorithm allows great flexibility, its mixing efficiency can vary hugely, and is poorly understood. Here we use mathematical analysis and simulation to explore the mixing efficiency of trans-model MCMC proposals, including the model-proposal probabilities and the proposal kernel for model parameters. Our analysis confirms the intuition that one should preferentially propose models with high posterior probabilities, and propose parameter values from the posterior as much as possible. Our results provide guidelines for constructing efficient trans-model MCMC algorithms. The principles are applied to MCMC algorithms in phylogenetic reconstruction using two real datasets for primates and mammals.

stat.CO

Using Variational Inference to Improve the Efficiency of MCMC Algorithms

Bayesian statistics makes inference based on Bayes' theorem, but the posterior distribution of unknown parameters is typically analytically intractable. To estimate the posterior, two widely used numerical approximation methods are Markov Chain Monte Carlo (MCMC) and variational inference (VI). MCMC methods produce asymptotically exact samples but are computationally intensive, while VI methods are faster and more scalable but may lack accuracy. This paper proposes combining MCMC and VI to construct algorithms that leverage the strengths of both. The first proposed algorithm uses Gaussian variational inference (GVI) with various covariance structures to derive a linear transformation matrix for Hamiltonian Monte Carlo (HMC). This method improves the efficiency of HMC, particularly in high-dimensional and complex target distributions. The second algorithm combines a VI-based generative model, the variational auto-encoder (VAE), with the Metropolis-Hastings (MH) sampler. The resulting VAE-MH sampler is efficient and effectively traverses the parameter space, outperforming standard MCMC methods in identifying all modes of multi-modal distributions.

stat.CO

A Novel Framework Using Variational Inference with Normalizing Flows to Train Transport Reversible Jump Proposals

We propose a unified framework that employs variational inference (VI) with (conditional) normalizing flows (NFs) to train both between-model and within-model proposals for reversible jump Markov chain Monte Carlo, enabling efficient trans-dimensional Bayesian inference. In contrast to the transport reversible jump (TRJ) of Davies et al. (2023), which optimizes forward KL divergence using pilot samples from the complex target distribution, our approach minimizes the reverse KL divergence, requiring only samples from a simple base distribution and largely reducing computational cost. Especially, we develop a novel trans-dimensional VI method with conditional NFs to fit the conditional transport proposal of Davies et al. (2023). We use RealNVP flows to learn the model-specific transport maps used for constructing proposals so that the calculation is parallelizable. Our framework also provides accurate estimates of marginal likelihoods, which may facilitate efficient model comparison and help design rejection-free proposals. Extensive numerical studies demonstrate that the TRJ method trained under our framework achieves faster mixing compared to existing baselines.

stat.ML

Efficient Mirror-type Kernels for the Metropolis-Hastings Algorithm

We propose a new Metropolis-Hastings (MH) kernel by introducing the Mirror move into the Metropolis adjusted Langevin algorithm (MALA). This new kernel uses the strength of one kernel to overcome the shortcoming of the other, and generates proposals that are distant from the current position, but still within the high-density region of the target distribution. The resulting algorithm can be much more efficient than both Mirror and MALA, while stays comparable in terms of computational cost. We demonstrate the advantages of the MirrorMALA kernel using a variety of one-dimensional and multi-dimensional examples. The Mirror and MirrorMALA are both special cases of the Mirror-type kernels, a new suite of efficient MH proposals. We use the Mirror-type kernels, together with a novel method of doing the whitening transformation on high-dimensional random variables, which was inspired by Tan and Nott, to analyse the Bayesian generalized linear mixed models (GLMMs), and obtain the per-time-unit efficiency that is 2--20 times higher than the HMC or NUTS algorithm.

stat.CO

Standardizing Type Ia supernovae using Near Infrared rebrightening time

Accurate standardisation of Type Ia supernovae (SNIa) is instrumental to the usage of SNIa as distance indicators. We analyse a homogeneous sample of 22 low-z SNIa, observed by the Carnegie Supernova Project (CSP) in the optical and near infra-red (NIR). We study the time of the second peak in the NIR band due to re-brightening, t2, as an alternative standardisation parameter of SNIa peak brightness. We use BAHAMAS, a Bayesian hierarchical model for SNIa cosmology, to determine the residual scatter in the Hubble diagram. We find that in the absence of a colour correction, t2 is a better standardisation parameter compared to stretch: t2 has a 1 sigma posterior interval for the Hubble residual scatter of [0.250, 0.257] , compared to [0.280, 0.287] when stretch (x1) alone is used. We demonstrate that when employed together with a colour correction, t2 and stretch lead to similar residual scatter. Using colour, stretch and t2 jointly as standardisation parameters does not result in any further reduction in scatter, suggesting that t2 carries redundant information with respect to stretch and colour. With a much larger SNIa NIR sample at higher redshift in the future, t2 could be a useful quantity to perform robustness checks of the standardisation procedure.

astro-ph.CO

A Corrected and More Efficient Suite of MCMC Samplers for the Multinomal Probit Model

The multinomial probit (MNP) model is a useful tool for describing discrete-choice data and there are a variety of methods for fitting the model. Among them, the algorithms provided by Imai and van Dyk (2005a), based on Marginal Data Augmentation, are widely used, because they are efficient in terms of convergence and allow the possibly improper prior distribution to be specified directly on identifiable parameters. Burgette and Nordheim (2012) modify a model and algorithm of Imai and van Dyk (2005a) to avoid an arbitrary choice that is often made to establish identifiability. There is an error in the algorithms of Imai and van Dyk (2005a), however, which affects both their algorithms and that of Burgette and Nordheim (2012). This error can alter the stationary distribution and the resulting fitted parameters as well as the efficiency of these algorithms. We propose a correction and use both a simulation study and a real-data analysis to illustrate the difference between the original and corrected algorithms, both in terms of their estimated posterior distributions and their convergence properties. In some cases, the effect on the stationary distribution can be substantial.

stat.CO

Metropolis-Hastings within Partially Collapsed Gibbs Samplers

The Partially Collapsed Gibbs (PCG) sampler offers a new strategy for improving the convergence of a Gibbs sampler. PCG achieves faster convergence by reducing the conditioning in some of the draws of its parent Gibbs sampler. Although this can significantly improve convergence, care must be taken to ensure that the stationary distribution is preserved. The conditional distributions sampled in a PCG sampler may be incompatible and permuting their order may upset the stationary distribution of the chain. Extra care must be taken when Metropolis-Hastings (MH) updates are used in some or all of the updates. Reducing the conditioning in an MH within Gibbs sampler can change the stationary distribution, even when the PCG sampler would work perfectly if MH were not used. In fact, a number of samplers of this sort that have been advocated in the literature do not actually have the target stationary distributions. In this article, we illustrate the challenges that may arise when using MH within a PCG sampler and develop a general strategy for using such updates while maintaining the desired stationary distribution. Theoretical arguments provide guidance when choosing between different MH within PCG sampling schemes. Finally we illustrate the MH within PCG sampler and its computational advantage using several examples from our applied work.

stat.CO