arXiv · 2509.19344
Performance of Large Language Models in Answering Critical Care Medicine Questions
Abstract
Large Language Models have been tested on medical student-level questions, but their performance in specialized fields like Critical Care Medicine (CCM) is less explored. This study evaluated Meta-Llama 3.1 models (8B and 70B parameters) on 871 CCM questions. Llama3.1:70B outperformed 8B by 30%, with 60% average accuracy. Performance varied across domains, highest in Research (68.4%) and lowest in Renal (47.9%), highlighting the need for broader future work to improve models across various subspecialty domains.
Explore related subjects
Keep this discovery
Mahmoud Alwakeel, Aditya Nagori, An-Kwok Ian Wong, Neal Chaisson, Vijay Krishnamoorthy, Rishikesan Kamaleswaran. 2025-09-16. Performance of Large Language Models in Answering Critical Care Medicine Questions. https://arxiv.org/abs/2509.19344
Cite the original work for its findings. Save a collection to share your selection of sources.