arXiv · 2503.17620
A Case Study of Scalable Content Annotation Using Multi-LLM Consensus and Human Review
Abstract
Content annotation at scale remains challenging, requiring substantial human expertise and effort. This paper presents a case study in code documentation analysis, where we explore the balance between automation efficiency and annotation accuracy. We present MCHR (Multi-LLM Consensus with Human Review), a novel semi-automated framework that enhances annotation scalability through the systematic integration of multiple LLMs and targeted human review. Our framework introduces a structured consensus-building mechanism among LLMs and an adaptive review protocol that strategically engages human expertise. Through our case study, we demonstrate that MCHR reduces annotation time by 32% to 100% compared to manual annotation while maintaining high accuracy (85.5% to 98%) across different difficulty levels, from basic binary classification to challenging open-set scenarios.
Explore related subjects
Keep this discovery
Mingyue Yuan, Jieshan Chen, Zhenchang Xing, Gelareh Mohammadi, Aaron Quigley. 2025-03-22. A Case Study of Scalable Content Annotation Using Multi-LLM Consensus and Human Review. https://arxiv.org/abs/2503.17620
Cite the original work for its findings. Save a collection to share your selection of sources.