arXiv · 2509.02855
IDEAlign: Comparing Ideas of Large Language Models to Domain Expert
Abstract
Large language models (LLMs) are increasingly used to produce open-ended, interpretive annotations, yet there is no validated, scalable measure of idea-level similarity to expert annotations. We (i) introduce the content evaluation of LLM annotations as a core, understudied task, (ii) propose IDEAlign for capturing expert similarity judgments via pick-the-odd-one-out tasks, and (iii) benchmark various similarity methods (text embeddings, topic models, and LLM-as-a-judge) against these human ratings. Applying this approach to two real-world educational datasets (e.g., interpreting math reasoning and feedback generation), we find that most metrics fail to capture the nuanced dimensions of similarity meaningful to experts. LLM-as-a-judge performs best (11~18% improvement over other methods) but still falls short of expert alignment, making it useful as a triage tool rather than a substitute for human review. Our work demonstrates the difficulty of evaluating open-ended LLM annotations at scale, and positions IDEAlign as a reusable protocol for benchmarking on this task to help guide responsible deployment of LLMs.
Explore related subjects
Keep this discovery
Hyunji Nam, Lucia Langlois, James Malamut, Mei Tan, Dorottya Demszky. 2025-09-02. IDEAlign: Comparing Ideas of Large Language Models to Domain Expert. https://arxiv.org/abs/2509.02855
Cite the original work for its findings. Save a collection to share your selection of sources.