arXiv · 2609.31719
Distributional Metrics for Evaluating Spoken Conversational Systems
Abstract
Evaluating conversational systems is a difficult and unresolved problem. We introduce the Conversational Distribution Score (CDS), which compares distributions of conversational behaviour using human conversations as a reference. CDS describes speech rate, syllabic rhythm, and turn interaction through eight interpretable features plus a separate two-feature semantic baseline. We compare conversations with two reference scales: one based on conversational success within human dialogue and another contrasting human and synthetic dialogue. Using listener judgments from out-of-domain goal--oriented dialogues, we examine system ranking, preferences between conversations, and ranking stability. Composite CDS recovers five of six listener system comparisons while individual features show strong correlation with listener preferences between conversations. We examine how many minutes and conversations are required before rankings stabilize. These findings support distributional comparisons as a complement to specific interactional metrics to evaluate conversations and conversational models while showing their interpretable value.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Shree Harsha Bokkahalli Satish, Erica Cooper, Patrícia Schmidtová, Maike Züfle, Éva Székely, Nicholas Sanders, Ondřej Klejch. 2026-09-21. Distributional Metrics for Evaluating Spoken Conversational Systems. https://arxiv.org/abs/2609.31719
Cite the original work for its findings. Save a collection to share your selection of sources.