arXiv · 2510.25577
Lost in Phonation: Voice Quality Variation as an Evaluation Dimension for Speech Foundation Models
Abstract
Recent advances in Speech Foundation Models (SFMs) enable direct processing of raw audio, allowing models to respond to subtle paralinguistic variation. However, how these models interpret non-lexical cues remains largely unstudied. We introduce VQ-Bench, a controlled evaluation suite featuring a parallel dataset of synthesized modal, breathy, creaky, and end-creak phonation types. We evaluate SFM sensitivity through open-ended generation across four ecologically valid domains, alongside speech emotion recognition. Our results reveal performance gaps: while a leading commercial API failed basic biometric sanity checks, other models exhibited systematic shifts in agency, empathy, and leadership based on phonation. Our findings also highlight gender asymmetries in salary and leadership endorsements, demonstrating that SFMs may mirror or amplify human social biases. This work establishes a reproducible framework for ensuring responsible paralinguistic interpretation in speech-based AI.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Harm Lameris, Shree Harsha Bokkahalli Satish, Joakim Gustafson, Éva Székely. 2025-10-29. Lost in Phonation: Voice Quality Variation as an Evaluation Dimension for Speech Foundation Models. https://arxiv.org/abs/2510.25577
Cite the original work for its findings. Save a collection to share your selection of sources.