Searcharxiv⌕ Search

arXiv subjects

Yunzhi Shen

Publications and source records attributed to Yunzhi Shen.

2 recordsLinked to original sources

Rubric-as-Experts: Case-Specific MQM Rubrics for Translation Error Span Detection

Large language models (LLMs) have shown potential for reference-free span-level translation quality estimation (QE), yet existing approaches based on Multidimensional Quality Metrics (MQM) typically rely on fixed rubric configurations shared across translation instances. However, translation instances often differ substantially in error complexity, ambiguity, and required evaluation granularity, making static rubric allocation suboptimal for span-level error detection. We find that larger MQM subtype spaces improve error coverage but also introduce more false positives, while different translation instances prefer different rubric granularities, suggesting that evaluation spaces should be allocated dynamically for each case. Therefore, we propose a case-specific dynamic rubric framework that adaptively constructs MQM evaluation spaces for individual translation instances. Unlike methods that generate fully free-form rubrics, our framework remains grounded in the predefined MQM taxonomy while dynamically selecting suitable subtype spaces and evaluation granularity for different cases. Experiments on span-level QE benchmarks from the Conference on Machine Translation (WMT) across multiple model scales demonstrate that the proposed framework consistently improves F1 and Matthews Correlation Coefficient (MCC).

cs.CL↗

PEGRL: Improving Machine Translation by Post-Editing Guided Reinforcement Learning

Reinforcement learning (RL) has shown strong promise for LLM-based machine translation, with recent methods such as GRPO demonstrating notable gains; nevertheless, translation-oriented RL remains challenged by noisy learning signals arising from Monte Carlo return estimation, as well as a large trajectory space that favors global exploration over fine-grained local optimization. We introduce \textbf{PEGRL}, a \textit{two-stage} RL framework that uses post-editing as an auxiliary task to stabilize training and guide overall optimization. At each iteration, translation outputs are sampled to construct post-editing inputs, allowing return estimation in the post-editing stage to benefit from conditioning on the current translation behavior, while jointly supporting both global exploration and fine-grained local optimization. A task-specific weighting scheme further balances the contributions of translation and post-editing objectives, yielding a biased yet more sample-efficient estimator. Experiments on English$\to$Finnish, English$\to$Turkish, and English$\leftrightarrow$Chinese show consistent gains over RL baselines, and for English$\to$Turkish, performance on COMET-KIWI is comparable to advanced LLM-based systems (DeepSeek-V3.2). Our code and a set of representative pretrained models are publicly available at \url{https://github.com/NJUNLP/peg-rl} and \url{https://huggingface.co/collections/DGME/pegrl}

cs.CL↗