arXiv · 2606.13381
H\"older++: Improving the Quality-Coherence Trade-off in Multimodal VAEs
Abstract
Existing approaches for multimodal variational autoencoders (VAEs) face a trade-off between generative quality and coherence-i.e., they struggle to generate realistic and diverse samples that, at the same time, are semantically consistent across modalities. A recent work shows that using a simple approximation to H\"older pooling as an aggregation method improves coherence over the SOTA MMVAE+, despite assuming a single shared representation across all modalities. Yet, it slightly compromises sample diversity. Inspired by this insight, we propose H\"older++, a novel multimodal VAE that improves the generative quality-coherence trade-off through: (i) the first implementation of H\"older pooling without any approximation for multimodal VAEs; (ii) an extended architecture that models distinct shared and private (i.e., modality-specific) representations (H\"older+); and (iii) hierarchical inference that further enhances the disentanglement between the shared and private representations (H\"older++). Our experiments corroborate that H\"older++ consistently improves the generative quality-coherence trade-off, yields more structured latent spaces, and learns shared representations that are informative for downstream tasks.
Explore related subjects
Keep this discovery
Huyen Vo, María Martínez-García, Isabel Valera. 2026-06-11. H\"older++: Improving the Quality-Coherence Trade-off in Multimodal VAEs. https://arxiv.org/abs/2606.13381
Cite the original work for its findings. Save a collection to share your selection of sources.