arXiv · 2608.08283
Do Evaluation Metrics Detect Errors in Classical Chinese to English Translations?
Abstract
Although large language models can translate some historical languages surprisingly well, their usefulness in digital humanities workflows is limited by the lack of reliable evaluation. We investigate whether existing automatic evaluation metrics developed for modern languages are reliable in this setting, using translation from Classical Chinese to English as a test case. We introduce a diagnostic framework based on minimal pairs capturing error types salient in scholarly use, probing both reference-based and reference-free metrics for error sensitivity and tolerance to valid variation. We find that all metrics exhibit blind spots, however MetricX24 performs best overall. Our findings highlight the need for more robust and interpretable metrics for historically and culturally distinct translation settings.
Explore related subjects
Keep this discovery
Osvaldo Quinjica, Eric Bennett, Xinchen Yang, Andrew Schonebaum, Marine Carpuat. 2026-08-08. Do Evaluation Metrics Detect Errors in Classical Chinese to English Translations?. https://arxiv.org/abs/2608.08283
Cite the original work for its findings. Save a collection to share your selection of sources.