Searcharxiv⌕ Search

arXiv · 2610.04541

Autonomous Structuring of Radiology Reports Across Modalities at Archive Scale Using an Open-Weight Large Language Model

Abstract

Purpose: To develop and evaluate an open-weight large language model (LLM) pipeline that converts an entire archive of free-text radiology reports into structured reports without human oversight. Materials and Methods: In this retrospective study, a pipeline with 150 hierarchically organized templates was developed at one center and tested at a second center on reports from 2010 to 2025. The open-weight model gpt-oss-120B selects the template in three constrained-decoding steps and fills it on one local graphics processing unit. Template selection was scored against expert labels on 914 randomly sampled reports of five modalities, structuring quality on 920 radiography and CT reports corrected field by field by five residents. The pipeline then processed the complete archive of the second center. Proportions are reported with Wilson 95% confidence intervals (CIs). Results: An optimal template set was selected for 74.4% of reports (680 of 914; 95% CI: 71.5%, 77.1%) and an appropriate set for 82.3% (752 of 914; 95% CI: 79.7%, 84.6%), 87.7% for single-region and 54.1% for multi-region reports. Macro semantic textual similarity between output and corrected reference was 0.95 for radiography and 0.97 for CT, residents left 88.7% of 24,638 fields unchanged, and unsupported content was flagged in 1.0% and 1.5% of reports. Of 2,186,982 archive reports, 96.5% received structured output, 2,401,544 structured reports, at 1,258 reports per hour on one graphics processing unit. Conclusion: An open-weight LLM pipeline structured a complete multimodality report archive without human oversight with high content fidelity. Multi-region reports remained the main source of template errors.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Friedrich Puttkammer, Fabian Drexel, Marlene Fritzsche, Era Stambollxhiu, Miriam Kumpf, Lena Schmitzer, Lea Schumann, Lina Xu, Johannes Moll, Jannik Lübberstedt, Zeineb Ben Chaaben, Anirudh Narayanan, Hartmut Häntze, Renato Cuocolo, Antonios Billis, Alexander Löser, Jawed Nawabi, Marcus R. Makowski, Cosmin I. Bercea, Shahrooz Faghihroohi, Lisa C. Adams, Keno K. Bressem. 2026-10-03. Autonomous Structuring of Radiology Reports Across Modalities at Archive Scale Using an Open-Weight Large Language Model. https://arxiv.org/abs/2610.04541

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Thought-Like-Pro: Enhancing Reasoning of Large Language Models through Self-Bootstrapped Prolog-based Chain-of-Thought

Large language models have demonstrated remarkable capabilities as general-purpose assistants, excelling in a wide range of reasoning tasks and supporting various aspects of daily web usage. This achievement represents a significant step toward achieving artificial general intelligence. Despite these advancements, the effectiveness of large language models often hinges on the specific prompting strategies employed, and there remains a lack of a robust framework to facilitate learning and generalization across diverse reasoning tasks. To address these challenges, we introduce a novel learning framework, Thought-Like-Pro. In this framework, we utilize imitation learning to imitate the Chain-of-Thought process which is verified and translated from reasoning trajectories generated by a symbolic Prolog logic engine. This framework proceeds in a prompt-guided but self-bootstrapped manner, that enables large language models to formulate rules and statements from given instructions and leverage the symbolic Prolog engine to derive results. Subsequently, large language models convert Prolog-derived successive reasoning trajectories into natural language chain-of-thought for imitation learning. The empirical findings indicate that our proposed approach greatly improves the reasoning capacity of large language models. By employing model averaging techniques, our method exhibits only a marginal decline in performance for distributional extrapolation tasks, showing robust generalization capabilities. We present a technical approach that integrates symbolic reasoning with language modeling, with the potential to support the development of large language models as cognitively inspired systems. The part of the dataset we used has been open-sourced.

cs.AI↗

AgentFly: Scaling Agentic Reinforcement Learning with Unified Resource System

Methods to build LLM agents have evolved from prompt engineering and supervised finetuning to agentic reinforcement learning (agentic RL). However, agentic RL remains bottlenecked by its surrounding systems: agents must interact with heterogeneous environments, such as sandboxes, model services, and external APIs. Their allocation, reuse, and lifecycle dominate rollout cost and cap the scale at which training becomes practical. In this work, we present AgentFly, an agentic RL framework built with a unified resource layer that treats each of these environments as a distinct, typed resource scheduled through one engine, with per-tool acquisition for multi-turn reuse, asynchronous backpressure, and rollout versus global-scoped lifecycles. AgentFly adopts a four-layer design: (I) agent layer that abstracts the agent, tool, and reward concepts, decomposing agentic RL into defining agents, tools, and reward functions; (II) rollout layer that composes these into agent loops and computes rewards; (III) context layer that organizes rollouts, injects contextual information, and arranges resources; and (IV) a low-level resource layer that performs resource management. We provide a suite of prebuilt tools and environments, demonstrate successful agent training across multiple tasks and models, and report the first controlled cross-framework throughput comparison against agentic RL frameworks.

cs.AI↗

Neural Architecture Discovery via Autonomous Evolution

Recent progress in LLM agents has advanced the prospect of autonomous research. Yet whether AI can complete difficult long-horizon tasks, especially those that advance AI research itself, remains largely unexplored. We present ASI-Arch, a system for AI-driven AI research that autonomously conducts neural architecture research through a closed-loop research-experiment-analyze-update process. Applied to linear attention, ASI-Arch ran 1,773 iterative experiments and discovered 105 state-of-the-art architectures. Its best architecture improves over DeltaNet by nearly three times the gain achieved by Mamba2. Beyond the final performance gains, we analyze the contributions of different parts of the framework in this hard research setting, shedding light on what enables autonomous progress in complex AI research tasks.

cs.AI↗