arXiv · 2607.16330
Local Brushstroke Quality Assessment via Vision-Language Feedback
Abstract
This paper investigates whether multimodal LLMs can evaluate local brushstroke quality in calligraphy and generate educationally useful natural language feedback. We construct an evaluation framework in which three multimodal LLMs (GPT-4o, Claude Sonnet 4, and Gemini 2.5 Flash) assess before-after image pairs of calligraphic works using a five-point ordinal scale, and compare their outputs against scores assigned by three expert calligraphers. We additionally examine a Retrieval-Augmented Generation (RAG) variant of Claude as a preliminary condition. Results show that all models achieve useful levels of absolute score accuracy (MAE), with GPT-4o performing best (MAE = 0.885). However, none of the models produce statistically significant overall rank correlations with human experts (Kendall's tau). Vocabulary analysis of generated rationales reveals characteristic evaluative biases in each model, and RAG is shown to improve rank correlation while worsening absolute accuracy, constituting an important negative result for text-based rule injection.
Explore related subjects
Keep this discovery
Mio Mitamura, Hirokatsu Kataoka. 2026-07-15. Local Brushstroke Quality Assessment via Vision-Language Feedback. https://arxiv.org/abs/2607.16330
Cite the original work for its findings. Save a collection to share your selection of sources.