Searcharxiv⌕ Search

arXiv · 2609.39525

Mitigating Representation Gaps in Amortized Bayesian Inference with Auxiliary Supervision

Abstract

Casting Bayesian inference as a neural network optimization problem targeting an amortized posterior is attractive, as it extends to otherwise intractable statistical models and offers near instantaneous inference for new datasets after prepaying the training cost. Although theory guarantees faithfulness under ideal convergence, practical amortized inference still requires iterating over architectures and optimization choices and ultimately ``satisficing'' under finite simulation, compute, and time budgets. Even the best-performing solution may thus retain avoidable representation gaps that typically require problem-specific fixes. Here, we propose a generic alternative which improves training dynamics with auxiliary guidance losses applied to internal representations. Specifically, we show how such guidance leads to faster convergence when training data is abundant and to better performance when it is scarce. We formalize representation gaps as getting stuck in a local optimum at the information bottleneck between the parts of the network tasked with feature learning and those tasked with conditional distribution learning, and offer a generic diagnostic to separate summary failures from inference failures. Finally, we demonstrate that auxiliary supervision improves convergence speed and accuracy on a range of challenging real-world inference problems.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Hans Olischläger, Svenja Jedhoff, Šimon Kucharský, Aayush Mishra, Stefan T. Radev, Paul Bürkner. 2026-09-30. Mitigating Representation Gaps in Amortized Bayesian Inference with Auxiliary Supervision. https://arxiv.org/abs/2609.39525

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Bringing Generative Learning to Representation Learning: Self-Supervised Transfer Learning as Distribution Matching

Most self-supervised learning objectives defend against collapse but leave the target representation law unspecified. We formulate representation learning as Distribution Matching (DM), learning an augmentation-invariant encoder whose induced law matches an explicit geometric reference. The reference law specifies what the learned representation distribution should look like, whereas a separately chosen discrepancy determines how deviations from this target are measured; here we use Mallows distance. The DM framework reveals a directional inverse: generative learning maps a tractable reference to data, whereas representation learning maps data to a designed reference law. We connect the population objective to class-centre separation and classification error and prove a non-asymptotic neural-sieve guarantee. Simulations and image benchmarks show manifold rectification, fine-grained structure and transfer across label spaces.

stat.ML↗

Conformalized Regression for Continuous Bounded Outcomes

Regression problems with continuous bounded outcomes frequently arise in statistical and machine learning applications, such as the analysis of rates and proportions. A central challenge in this setting is predicting the response at a new covariate value. Most of the existing literature has focused either on point prediction or on interval prediction based on asymptotic approximations. We develop conformal prediction intervals for bounded outcomes within the framework of transformation regression models, encompassing widely used models such as beta regression and logit-normal regression. We construct non-conformity scores based on model-aligned residuals and identify a quantile-residual score that is particularly well suited to bounded outcomes, bridging normalized conformal prediction and distributional conformal prediction. This score accounts for both the heteroscedasticity inherent in such data and the asymmetry that emerges near the boundaries of the response space. We establish marginal validity and asymptotic conditional validity for both full and split conformal prediction, holding under model misspecification. A comprehensive simulation study confirms that both methods empirically attain valid finite-sample coverage, including cases under model misspecification. A real-data application demonstrates their practical performance against bootstrap-based alternatives.

stat.ML↗

Beyond Uncertainty Sets: Leveraging Optimal Transport to Extend Conformal Predictive Distributions to Multivariate Settings

Conformal prediction (CP) constructs uncertainty sets for model outputs with finite-sample coverage guarantees. Yet ranking scores is straightforward only when they are scalar-valued, limiting CP to real-valued scores or ad-hoc one-dimensional reductions. Vector-valued scores arise naturally in multi-output regression and model aggregation, where each predictor in an ensemble provides its own score. Optimal transport (OT) defines vector ranks and center-outward multivariate quantile regions, though generally with asymptotic coverage guarantees. Applying a fixed transport map learned from calibration data to a new point introduces an uncontrolled approximation error. We restore finite-sample, distribution-free coverage by conformalizing vector-valued OT quantile regions. Each candidate's rank is defined by transporting the calibration scores augmented with that candidate's score, preserving the symmetry needed for validity. This appears to require a continuum of OT problems. However, we prove that the optimal assignment is piecewise constant across a fixed polyhedral partition of score space. This lets us characterize the entire prediction set in $O(n^3)$ time, matching the cost of a single assignment solve. It also addresses a limitation of prediction sets: they indicate which outcomes are plausible, but not their relative likelihood. In one dimension, conformal predictive distributions (CPDs) fill this gap by producing a predictive distribution with finite-sample calibration. Extending CPDs beyond one dimension remained an open problem. We construct, to our knowledge, the first multivariate CPDs with finite-sample calibration: a center-outward predictive distribution whose derived uncertainty regions have conformal coverage. We present both conservative and exact randomized versions; the latter generalizes the classical Dempster-Hill procedure.

stat.ML↗