arXiv · 2601.03321
HERO: Hierarchical Evidential Reasoning Optimization for Radiology Report Generation via Reason-then-Summarize
Abstract
Multimodal Large Language Models (MLLMs) have substantially advanced Radiology Report Generation (RRG), yet aligning them through reinforcement learning (RL) remains challenging due to heterogeneous medical supervision. Vanilla Group Relative Policy Optimization (GRPO) assigns uniform credit across the entire generation, leading to segment interference, token dilution, and evidence--diagnosis decoupling, which exacerbates clinical hallucinations. We propose HERO (Hierarchical Evidential Reasoning Optimization), a factorized policy optimization framework that aligns heterogeneous supervision with three optimization granularities. HERO separately optimizes reasoning, diagnosis, and evidence grounding through complementary segment-, token-, and completion-level optimization with a heterogeneous reward formulation covering diagnostic accuracy, reasoning quality, and think--answer consistency. Experiments on MIMIC-CXR and IU-Xray show that HERO outperforms strong supervised and reinforcement learning baselines, achieving state-of-the-art clinical efficacy while producing more evidence-grounded and think--answer-consistent reports, thereby substantially mitigating clinical hallucinations.
Explore related subjects
Keep this discovery
Kun Zhao, Guodong Liu, Hui Ji, Siyuan Dai, Pan Wang, Jifeng Song, Chenghua Lin, Liang Zhan, Haoteng Tang. 2026-01-06. HERO: Hierarchical Evidential Reasoning Optimization for Radiology Report Generation via Reason-then-Summarize. https://arxiv.org/abs/2601.03321
Cite the original work for its findings. Save a collection to share your selection of sources.