arXiv · 2511.15145
Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding
Abstract
Human voice encodes both identity and paralinguistic cues, yet encoders in large audio-language models (LALMs) rarely balance both aspects. In this work, we present a study toward building a general-purpose voice encoder that captures nuanced voice cues. Through a comprehensive evaluation, we find that multi-task training yields the most balanced representations, whereas contrastive language-audio pretraining (CLAP) primarily improves retrieval without enhancing paralinguistic understanding. Our final encoder, Auden-Voice, also demonstrates strong performance when integrated with LLMs. The code and training recipes will be released with the audio understanding toolkit Auden.
Explore related subjects
Keep this discovery
Mingyue Huo, Wei-Cheng Tseng, Yiwen Shao, Hao Zhang, Dong Yu. 2025-11-19. Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding. https://arxiv.org/abs/2511.15145
Cite the original work for its findings. Save a collection to share your selection of sources.