Searcharxiv⌕ Search

arXiv · 2610.01448

On skew-symmetric distributions and their use in Monte Carlo sampling algorithms: coordinate-free, Gibbs-style and manifold versions of the Barker proposal

Abstract

Skew-symmetric probability distributions provide a principled mechanism for incorporating gradient information into Markov chain Monte Carlo algorithms. Here we review the (preconditioned) Barker proposal, a Metropolis--Hastings algorithm built on skew-symmetric distributions, and motivate its design. We then introduce three natural extensions. First, we propose coordinate-free variants of the Barker algorithm. Second, we introduce a Gibbs-style Barker algorithm that re-evaluates the gradient at each partially updated coordinate. Third, we derive a simplified manifold Barker algorithm, producing a manifold sampler with enhanced robustness compared to natural comparators. Numerical experiments demonstrate that the Gibbs-style variant improves raw sampling efficiency on correlated targets, that the coordinate-free variants offer limited practical advantage over the standard Barker proposal once computational costs are accounted for, and that the simplified manifold Barker algorithm can achieve significant advantages over simplified manifold MALA when the local geometric structure of the target is irregular or unreliable.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Minh Vu, Samuel Livingstone, Pantelis Samartsidis. 2026-10-01. On skew-symmetric distributions and their use in Monte Carlo sampling algorithms: coordinate-free, Gibbs-style and manifold versions of the Barker proposal. https://arxiv.org/abs/2610.01448

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Efficient Solvers for SLOPE in R, Python, Julia, and C++

We present a suite of packages in R, Python, Julia, and C++ that efficiently solve the Sorted L-One Penalized Estimation (SLOPE) problem. The packages feature a highly efficient hybrid coordinate descent algorithm that fits generalized linear models (GLMs) and supports a variety of loss functions, including Gaussian, binomial, Poisson, and multinomial logistic regression. Our implementation is designed to be fast, memory-efficient, and flexible. The packages support a variety of data structures (dense, sparse, and out-of-memory matrices) and are designed to efficiently fit the full SLOPE path as well as handle cross-validation of SLOPE models, including the relaxed SLOPE. We present examples of how to use the packages and benchmarks that demonstrate the performance of the packages on both real and simulated data and show that our packages outperform existing implementations of SLOPE in terms of speed.

stat.CO↗

SSLfmm: An R Package for Semi-Supervised Learning with Mixed Missingness

Partially labelled samples arise when features are observed for data, but class labels are available for only a subset. In such settings, the mechanism governing label availability may itself contain information relevant to classification, yet it is typically left unmodelled in standard semi-supervised learning procedures. The SSLfmm package implements likelihood-based Gaussian finite-mixture classification in which the label-missingness process is modelled jointly with the class distribution. It supports complete-case, missing completely at random (MCAR), entropy-based missing at random (MAR), and mixed analyses in which MCAR and MAR mechanisms may both contribute. For the mixed mechanism, the source of a missing label may be known or unknown. A common R interface is provided for model fitting, prediction, performance assessment, and simulation. We describe the statistical formulation and software implementation and demonstrate its use through reproducible simulation and a semi-synthetic Blood Transfusion application.

stat.CO↗

Group recovery after trimming, and level-free flagging, in robust clusterwise regression

Trimming methods for robust clusterwise regression discard a fixed fraction of the data. Too low a level breaks the fit; too generous a level can trim away a small group. We first propose a group-recovery step that can follow any trimming or flagging method: it searches the discarded units for a line, tests whether those near it form a peak rather than a band, and restores the line as a group when the likelihood of Gaussian groups plus uniform noise improves by a margin like that of the Bayesian information criterion. In simulations it repaired the failures of a generous TCLUST-REG level with unequal groups, but not with three or four groups. The second proposal, ESF (exact-subsample flagging), is a flagging procedure without a trimming level: it solves the clusterwise least-squares problem exactly on small subsamples, flags units far from the best fit, and draws later subsamples from the rest. Two constants stand in for the level: a subsample size, set from a lower bound on the smallest group proportion, and a cap on the flagged set. It is meant for data about whose contamination nothing is known: TCLUST-REG at the fixed level 0.30, followed by reweighting and the recovery step, was as accurate as ESF on average up to a fifth of outliers, and higher levels were more accurate beyond. The flagged fraction estimates the contamination only when the errors are close to Gaussian. On taxi fares with known tariffs, ESF found both in every sample.

stat.CO↗