arXiv · 2210.11835
A Textless Metric for Speech-to-Speech Comparison
Abstract
In this paper, we introduce a new and simple method for comparing speech utterances without relying on text transcripts. Our speech-to-speech comparison metric utilizes state-of-the-art speech2unit encoders like HuBERT to convert speech utterances into discrete acoustic units. We then propose a simple and easily replicable neural architecture that learns a speech-based metric that closely corresponds to its text-based counterpart. This textless metric has numerous potential applications, including evaluating speech-to-speech translation for oral languages, languages without dependable ASR systems, or to avoid the need for ASR transcription altogether. This paper also shows that for speech-to-speech translation evaluation, ASR-BLEU (which consists in automatically transcribing both speech hypothesis and reference and compute sentence-level BLEU between transcripts) is a poor proxy to real text-BLEU even when ASR system is strong.
Explore related subjects
Keep this discovery
Laurent Besacier, Swen Ribeiro, Olivier Galibert, Ioan Calapodescu. 2022-10-21. A Textless Metric for Speech-to-Speech Comparison. https://arxiv.org/abs/2210.11835
Cite the original work for its findings. Save a collection to share your selection of sources.