arXiv · 2507.18061
TELEVAL: A Benchmark Designed for Spoken Language Models in Chinese Interactive Scenarios
Abstract
Spoken Language Models (SLMs) are expected to support natural spoken interaction beyond task completion. However, existing SLM benchmarks primarily evaluate semantic correctness in structured settings and provide limited assessment of interactional behavior grounded in acoustic context. To address this gap, we introduce TELEVAL, a large-scale SLM benchmark for Chinese spoken interaction in instruction-free, audio-conditioned settings. TELEVAL evaluates two complementary aspects: (1) Reliable Content Fulfillment, which measures semantic accuracy of SLMs under diverse acoustic and linguistic conditions, and (2) Interactional Appropriateness, which assesses whether models produce natural and appropriate responses by implicitly grounding behavior in auditory cues. Experiments show that while models perform competitively on semantic tasks, their performance degrades under acoustic variability and in interactional settings. We observe consistent degradation from perceptual instability to interactional errors, and further identify a recurring failure pattern, termed the "Caption Trap", where models tend to describe perceived audio signals rather than produce appropriate interactive responses. These results indicate that current SLMs remain insufficiently aligned with the requirements of natural spoken interaction. TELEVAL provides a targeted framework for evaluating and analyzing interactional behavior in SLMs.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Zehan Li, Hongjie Chen, Qing Wang, Yuxin Zhang, Jing Zhou, Hang Lv, Mengjie Du, Yaodong Song, Jie Lian, Jian Kang, Jie Li, Yongxiang Li. 2025-07-24. TELEVAL: A Benchmark Designed for Spoken Language Models in Chinese Interactive Scenarios. https://arxiv.org/abs/2507.18061
Cite the original work for its findings. Save a collection to share your selection of sources.