Rubric-as-Experts: Case-Specific MQM Rubrics for Translation Error Span Detection
Large language models (LLMs) have shown potential for reference-free span-level translation quality estimation (QE), yet existing approaches based on Multidimensional Quality Metrics (MQM) typically rely on fixed rubric configurations shared across translation instances. However, translation instances often differ substantially in error complexity, ambiguity, and required evaluation granularity, making static rubric allocation suboptimal for span-level error detection. We find that larger MQM subtype spaces improve error coverage but also introduce more false positives, while different translation instances prefer different rubric granularities, suggesting that evaluation spaces should be allocated dynamically for each case. Therefore, we propose a case-specific dynamic rubric framework that adaptively constructs MQM evaluation spaces for individual translation instances. Unlike methods that generate fully free-form rubrics, our framework remains grounded in the predefined MQM taxonomy while dynamically selecting suitable subtype spaces and evaluation granularity for different cases. Experiments on span-level QE benchmarks from the Conference on Machine Translation (WMT) across multiple model scales demonstrate that the proposed framework consistently improves F1 and Matthews Correlation Coefficient (MCC).