arXiv · 2609.10466
Phoneme-Aware Pronunciation Representations for L2-English L1-Background Accent Identification
Abstract
We study speaker-disjoint accent identification for L2 English, where the goal is to predict a speaker's first-language (L1) background from English pronunciation. Most existing systems classify accents using a single utterance-level representation, but such global representations can obscure pronunciation cues that depend on specific English phonemes. We propose a transcript-assisted model that makes phoneme information explicit during accent identification. Instead of representing an utterance only as a global speech embedding, we represent it as a sequence of pronunciation units, each combining acoustic evidence from a spoken segment with the aligned English phoneme for that segment. A frozen speech encoder provides the acoustic features, while the transcript is used only to obtain phoneme-level forced alignments. No word-level or sentence-level text representation is passed to the accent classifier. Under a four-fold speaker-disjoint protocol on L2-ARCTIC, our model achieves 81.41% accuracy and 81.21% macro-F1, the highest mean performance among the evaluated systems. Diagnostic ablations support the importance of phoneme-aligned token construction, while a Whisper-based ablation shows an additional gain from phoneme information.
Explore related subjects
Keep this discovery
Yangyang Qu, Massimiliano Todisco, Nicholas Evans. 2026-09-09. Phoneme-Aware Pronunciation Representations for L2-English L1-Background Accent Identification. https://arxiv.org/abs/2609.10466
Cite the original work for its findings. Save a collection to share your selection of sources.