arXiv · 2610.01082
Precision over Scale: A Polish-Silesian Benchmark and a Translation System Outperforming Open-Source and Commercial Models
Abstract
Dialectal machine translation remains challenging due to limited data and strong linguistic variation not captured by standard benchmarks, which often assume standardized and well-edited text. We study Polish-Silesian MT using neural and rule-based systems, evaluating on SiLTT - a new Pol-Szl testset, alongside established BOUQuET and FLORES benchmarks. Results show our rule-based system is consistently strongest on SiLTT and BOUQuET datasets and that TranslateGemma fine-tuned on a curated dataset improves over strong neural baselines but does not surpass the rule-based system in dialectal settings. We release SiLTT and our best neural model to support further research.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Grzegorz Kulik, Mikołaj Pokrywka, Adam Jatowt, Wojciech Kusa. 2026-10-01. Precision over Scale: A Polish-Silesian Benchmark and a Translation System Outperforming Open-Source and Commercial Models. https://arxiv.org/abs/2610.01082
Cite the original work for its findings. Save a collection to share your selection of sources.