arXiv · 2602.15778
*-PLUIE: Personalisable metric with Llm Used for Improved Evaluation
Abstract
Evaluating the quality of automatically generated text often relies on LLM-as-a-judge (LLM-judge) methods. While effective, these approaches are computationally expensive and require post-processing. To address these limitations, we build upon ParaPLUIE, a perplexity-based LLM-judge metric that estimates confidence over ``Yes/No'' answers without generating text. We introduce *-PLUIE, task specific prompting variants of ParaPLUIE and evaluate their alignment with human judgement. Our experiments show that personalised *-PLUIE achieves stronger correlations with human ratings while maintaining low computational cost.
Explore related subjects
Keep this discovery
Quentin Lemesle, Léane Jourdan, Daisy Munson, Pierre Alain, Jonathan Chevelu, Arnaud Delhay, Damien Lolive. 2026-02-17. *-PLUIE: Personalisable metric with Llm Used for Improved Evaluation. https://arxiv.org/abs/2602.15778
Cite the original work for its findings. Save a collection to share your selection of sources.