Searcharxiv⌕ Search

arXiv · 2610.09589

Dual- versus Single-Suggestion AI Support for Radiographic Interpretation in Residents: Randomized Multireader Study

Abstract

Purpose: To compare dual- and single-suggestion AI support for radiographic interpretation by residents, particularly when the shared AI suggestion was incorrect. Materials and Methods: This prospective, multicenter, randomized three-arm reader study was conducted at three hospitals in China from July to September 2026 (ChiCTR2600129243). After specialty stratification, 132 residents with fewer than 3 years of clinical experience were randomized 1:1:1 to GPT-5.4 alone (group A), GPT-5.4 plus Kimi-K2.6 (group B), or GPT-5.4 plus Gemini-3.6 Flash (group C); 123 were analyzed. Participants interpreted 60 radiographs before and after AI support. The primary outcome was accuracy change. Welch ANOVA and Holm-adjusted t tests compared support conditions; HC3 linear models assessed specialty interaction. Results: Among 123 residents (mean age, 24.1 years +/- 1.4; 65 women), radiology residents showed greater accuracy improvement with dual- than single-suggestion support (B-A, 6.69 percentage points [95% CI, 0.97-12.40]; C-A, 7.87 percentage points [95% CI, 1.64-14.11]; Holm-adjusted P = .030 for both), whereas accuracy change did not differ in non-radiology residents (P = .20). When GPT-5.4 was incorrect, AI-assisted accuracy was higher with dual- than single-suggestion support in radiology residents (40.1% and 40.4% vs 20.0%) and non-radiology residents (31.3% and 31.0% vs 12.1%) (all Holm-adjusted P < .001). The dual-suggestion effect differed by specialty (interaction difference, 10.44 percentage points; 95% CI, 4.36-16.52; P < .001). Conclusion: Dual-suggestion support may mitigate the influence of erroneous AI suggestions, with greater accuracy improvement observed in radiology but not non-radiology residents.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Lin Wu, Zhe Xu, Hongyi Wang, Feifei Zhou, Wei Deng, Chunlong Zhang, Yuting Zhu, Kaixiao Chen, Xiao Liang, Chen Yang, Yeyuan Chen, Hao Chen, Fuqing Zhou. 2026-10-07. Dual- versus Single-Suggestion AI Support for Radiographic Interpretation in Residents: Randomized Multireader Study. https://arxiv.org/abs/2610.09589

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Thought-Like-Pro: Enhancing Reasoning of Large Language Models through Self-Bootstrapped Prolog-based Chain-of-Thought

Large language models have demonstrated remarkable capabilities as general-purpose assistants, excelling in a wide range of reasoning tasks and supporting various aspects of daily web usage. This achievement represents a significant step toward achieving artificial general intelligence. Despite these advancements, the effectiveness of large language models often hinges on the specific prompting strategies employed, and there remains a lack of a robust framework to facilitate learning and generalization across diverse reasoning tasks. To address these challenges, we introduce a novel learning framework, Thought-Like-Pro. In this framework, we utilize imitation learning to imitate the Chain-of-Thought process which is verified and translated from reasoning trajectories generated by a symbolic Prolog logic engine. This framework proceeds in a prompt-guided but self-bootstrapped manner, that enables large language models to formulate rules and statements from given instructions and leverage the symbolic Prolog engine to derive results. Subsequently, large language models convert Prolog-derived successive reasoning trajectories into natural language chain-of-thought for imitation learning. The empirical findings indicate that our proposed approach greatly improves the reasoning capacity of large language models. By employing model averaging techniques, our method exhibits only a marginal decline in performance for distributional extrapolation tasks, showing robust generalization capabilities. We present a technical approach that integrates symbolic reasoning with language modeling, with the potential to support the development of large language models as cognitively inspired systems. The part of the dataset we used has been open-sourced.

cs.AI↗

MAPF-World: Action World Model for Multi-Agent Path Finding

Multi-agent path finding (MAPF) studies the problem of planning conflict-free paths for multiple agents from given start locations to designated goals, with applications in robot-assisted logistics and social navigation. Recent decentralized learned solvers have shown promise for large-scale MAPF, particularly when leveraging foundation models and large datasets. However, most existing methods rely on reactive policies, often resulting in congestion, deadlocks, and degraded generalization in high agent-density environments. To address these limitations, we propose MAPF-World, an autoregressive action world model for MAPF that unifies short-horizon local future prediction and action generation, enabling decision-making beyond immediate local observations. MAPF-World models short-horizon local dynamics by predicting the next local observation and neighboring agents' action intentions, capturing both spatial structures and temporal interaction patterns. We further introduce a spatio-agent positional encoding that integrates spatial awareness with agent-level semantics in Transformer-based architectures, facilitating more coordinated multi-agent behaviors. In addition, we augment existing MAPF benchmarks by introducing an automated map generator grounded in real-world urban layouts, aiming to narrow the gap between synthetic simulation and practical deployment scenarios. Extensive experiments across diverse map types and interaction settings demonstrate that MAPF-World achieves strong performance compared with existing learned solvers. Notably, it exhibits robust zero-shot generalization and maintains a high success rate even as agent density increases.

cs.AI↗

More Than Meets the Eye? Uncovering the Reasoning-Planning Disconnect in Training Vision-Language Driving Models

Vision-Language Model (VLM) driving agents promise explainable end-to-end autonomy by first producing natural-language reasoning and then predicting trajectory planning. However, whether planning is causally driven by this reasoning remains a critical but unverified assumption. To investigate this, we build DriveMind, a large-scale driving Visual Question Answering corpus with plan-aligned Chain-of-Thought (CoT), automatically generated from nuPlan. Our data generation process converts sensors and annotations into structured inputs and, crucially, separates priors from to-be-reasoned signals, enabling clean information ablations. Using DriveMind, we train representative VLM agents with Supervised Fine-Tuning and Group Relative Policy Optimization and evaluate them with nuPlan's metrics. Our results, unfortunately, indicate a consistent causal disconnect in reasoning-planning: removing ego/navigation priors causes large drops in planning scores, whereas removing CoT produces only minor changes. Attention analysis further shows that planning primarily focuses on priors rather than the CoT. Based on this evidence, we propose the Reasoning-Planning Decoupling Hypothesis, positing that the training-yielded reasoning is an ancillary byproduct rather than a causal mediator. To enable efficient diagnosis, we introduce a novel, training-free probe that measures an agent's reliance on priors by evaluating its planning robustness against minor input perturbations. In summary, we provide the community with a new dataset and a diagnostic tool to evaluate the causal fidelity of future models.

cs.AI↗