arXiv · 2607.05058
Context-Aware ASR for Mandarin Technical Lectures
Abstract
Technical lectures mix Mandarin speech with English technical terms. These terms carry the core meaning of the lecture, yet they occupy few characters. Character error rate (CER) therefore hides their recognition failures. We study whether lecture context helps recognize these terms. We build a term-rich Mandarin AI/ML lecture benchmark, and we define term-centric metrics that measure technical-term recognition directly. We then propose a two-pass, reference-free decoding method. The first pass runs segment-only ASR. We extract the most frequent technical terms from the first-pass hypotheses, and we prompt the recognizer with this self-built glossary in the second pass. Across five ASR backbones, the first-pass glossary raises term recall for every model and holds or lowers CER on all five. On Breeze-ASR-25 it lifts term recall from 52.50% to 60.13% while lowering CER, and a hybrid that adds a small external term list reaches 62.05% recall and 82.73% term precision. Lecture context, recovered from the model's own output, is a practical signal for technical-term recognition. Term-centric evaluation exposes errors that CER misses.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ho-Lam Chung, Yiming Chen, Hung-yi Lee. 2026-07-06. Context-Aware ASR for Mandarin Technical Lectures. https://arxiv.org/abs/2607.05058
Cite the original work for its findings. Save a collection to share your selection of sources.