arXiv · 2510.21317
Are These Even Words? Quantifying the Gibberishness of Generative Speech Models
Abstract
Significant research efforts are currently being dedicated to non-intrusive quality and intelligibility assessment, especially given how it enables curation of large scale datasets of in-the-wild speech data. However, with the increasing capabilities of generative models to synthesize high quality speech, new types of artifacts become relevant, such as generative hallucinations. While intrusive metrics are able to spot such sort of discrepancies from a reference signal, it is not clear how current non-intrusive methods react to high-quality phoneme confusions or, more extremely, gibberish speech. In this paper we explore how to factor in this aspect under a fully unsupervised setting by leveraging language models. Additionally, we publish a dataset of high-quality synthesized gibberish speech for further development of measures to assess implausible sentences in spoken language, alongside code for calculating scores from a variety of speech language models.
Explore related subjects
Keep this discovery
Danilo de Oliveira, Tal Peer, Jonas Rochdi, Timo Gerkmann. 2025-10-24. Are These Even Words? Quantifying the Gibberishness of Generative Speech Models. https://arxiv.org/abs/2510.21317
Cite the original work for its findings. Save a collection to share your selection of sources.