arXiv · 2609.10153
YallaMorph: A Benchmark for Evaluating Arabic Morphological Generation in Large Language Models
Abstract
Arabic morphology remains challenging for large language models, since fluent generation does not guarantee accurate morphosyntactic control. Existing Arabic evaluations mainly target downstream tasks and do not directly test controlled morphological generation from explicit lexical and feature-based input. We introduce YallaMorph, a large-scale benchmark for Arabic morphological generation covering verbs, nouns, adjectives, their cliticized forms, and invalid configurations. We evaluate multilingual and Arabic-oriented LLMs under diacritized and undiacritized settings over 600K benchmark entries. Results show that Arabic morphological generation remains difficult, especially for cliticized, unseen, and morphologically rare forms.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Mahmoud Reda, Salam Khalifa, Reham Marzouk, Nizar Habash. 2026-09-09. YallaMorph: A Benchmark for Evaluating Arabic Morphological Generation in Large Language Models. https://arxiv.org/abs/2609.10153
Cite the original work for its findings. Save a collection to share your selection of sources.