arXiv · 2609.37322
Mubric: Mutation Testing-Guided Rubric Generation for LLM Evaluation
Abstract
Rubric-based evaluation is widely used to assess LLM-based systems by decomposing response quality into task-specific scoring criteria. However, automatically generating rubrics that reliably capture task-specific quality requirements remains challenging. We introduce Mubric, a mutation testing-guided approach to rubric generation. Mutation testing, a classic software testing methodology, evaluates a test suite by injecting faults into programs and checking whether the tests detect them. We draw an analogy between test suites and rubrics: if a rubric captures an important quality requirement, introducing a corresponding defect into an otherwise high-quality response should reduce its score. Mubric first mines common defects from real pairs of preferred and dispreferred responses and abstracts these defects into reusable mutation operators, each specifying how to introduce a particular type of response defect. For a new task, it applies relevant operators to a reference response, checks whether the injected defects reduce response quality, and uses insufficiently penalized defects to refine the rubric. We evaluate Mubric on 703 tasks across four representative domains against six advanced rubric generation methods. Mubric achieves the highest overall evaluation accuracy, outperforming the strongest baseline by 7.48 percentage points.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jiayuxuan Yang, Jie M. Zhang, Yiling Lou, Zhenpeng Chen. 2026-09-29. Mubric: Mutation Testing-Guided Rubric Generation for LLM Evaluation. https://arxiv.org/abs/2609.37322
Cite the original work for its findings. Save a collection to share your selection of sources.