arXiv · 2609.35017
TermJudge: A Document-Level Metric Judging, Not Counting, Terminology in Machine Translation Evaluation
Abstract
Existing automatic metrics for evaluating terminological use in machine translation (MT) penalise any divergence from a fixed reference, conflating translation errors with the valid terminological variation that human translators routinely produce. We introduce TermJudge, a document-level terminology metric that assigns an interpretable verdict to every term occurrence: glossary-conforming occurrences are settled deterministically, while divergences are assessed under a two-step LLM-as-judge procedure using the full document context: the first detects and labels terminology errors; the second sorts valid document-level variations from inconsistencies. Validated against expert error annotations and document-level human MQM scores, TermJudge ranks first in both system- and segment-level meta-evaluation, ahead of glossary-conformity and quality-estimation baselines. When applied to eight systems translating academic documents, under two prompting conditions, we observe that glossary injection improves terminology translation in all paired comparisons, by removing genuine errors rather than valid variation. TermJudge is released as open-source code.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Nicolas Dahan, Fran{\cc}ois Yvon, Rachel Bawden. 2026-09-28. TermJudge: A Document-Level Metric Judging, Not Counting, Terminology in Machine Translation Evaluation. https://arxiv.org/abs/2609.35017
Cite the original work for its findings. Save a collection to share your selection of sources.