arXiv · 2609.34147
SPEAR-Gen: Generation-Aware Pre-training for Unified Speech Representations
Abstract
Speech understanding and generation place different demands on speech representations, and existing models are typically optimised towards one capability or the other. To reduce this gap, we introduce SPEAR-Gen, a speech representation model that learns a single representation for both capabilities. Task-aligned feature aggregation consolidates complementary linguistic and paralinguistic information across a frozen encoder into discrete targets for masked prediction, while a coarse-to-fine objective combines log-Mel reconstruction with residual flow matching to preserve spectral structure and fine-grained acoustic variation. Experiments on SUPERB and speech resynthesis show that SPEAR-Gen maintains strong understanding performance while substantially improving resynthesis quality and speaker preservation. These results demonstrate that a single speech representation can effectively support both understanding and generation.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Xiaoyu Yang, Arthur Hinsvark, Antonios Alexos, Osama Hanna, Philip C. Woodland, Yiting Lu. 2026-09-28. SPEAR-Gen: Generation-Aware Pre-training for Unified Speech Representations. https://arxiv.org/abs/2609.34147
Cite the original work for its findings. Save a collection to share your selection of sources.