arXiv · 2605.03297
Contrastive Regularization for Accent-Robust ASR
Abstract
ASR systems based on self-supervised acoustic pretraining and CTC fine-tuning achieve strong performance on native speech but remain sensitive to accent variability. We investigate supervised contrastive learning (SupCon) as a lightweight, accent-invariant auxiliary objective for CTC fine-tuning. An utterance-level contrastive loss regularizes encoder representations without architectural modification or explicit accent supervision. Experiments on the L2-ARCTIC benchmark show consistent WER reductions across multiple pretrained encoders, with up to 25 -- 29\% relative reduction under unseen-accent evaluation. Analysis using within-transcript cosine dispersion indicates that SupCon promotes more compact and stable representation geometry under accent variability. Overall, SupCon provides an effective and model-agnostic regularization strategy for improving accent robustness.
Explore related subjects
Keep this discovery
Van-Phat Thai, Aradhya Dhruv, Duc-Thinh Pham, Sameer Alam. 2026-05-05. Contrastive Regularization for Accent-Robust ASR. https://arxiv.org/abs/2605.03297
Cite the original work for its findings. Save a collection to share your selection of sources.