TY - RPRT TI - Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment AU - I. F. Atasoy AU - B. Mutlu AU - E. A. Sezer AU - A. Wahdan PY - 2026 UR - https://arxiv.org/abs/2605.08462 ID - 2605.08462 ER -