arXiv · 2605.06276
Linear Semantic Segmentation for Low-Resource Spoken Dialects
Abstract
Semantic segmentation is a core component of discourse analysis, yet existing models are primarily developed and evaluated on high-resource written text, limiting their effectiveness on low-resource spoken varieties. In particular, dialectal Arabic exhibits informal syntax, code-switching, and weakly marked discourse structure that challenge standard segmentation approaches. In this paper, we introduce a new multi-genre benchmark (more than 1000 samples) for semantic segmentation in conversational Arabic, focusing on dialectal discourse. The benchmark covers transcribed casual telephone conversations, code-switched podcasts, broadcast news, and expressive dialogue from novels, and was annotated and validated by native Arabic annotators. Using this benchmark, we show that segmentation models performing well on MSA news genres degrade on dialectal transcribed speech. We further propose a segmentation model that targets local semantic coherence and robustness to discourse discontinuities, consistently outperforming strong baselines on dialectal non-news genres. The benchmark and approach generalize to other low-resource spoken languages.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Kirill Chirkunov, Younes Samih, Abed Alhakim Freihat, Hanan Aldarmaki. 2026-05-07. Linear Semantic Segmentation for Low-Resource Spoken Dialects. https://arxiv.org/abs/2605.06276
Cite the original work for its findings. Save a collection to share your selection of sources.