arXiv · 2510.04338
Evaluation of Clinical Trials Reporting Quality using Large Language Models
Abstract
Reporting quality is an important topic in clinical trial research articles, as it can impact clinical decisions. In this article, we test the ability of large language models to assess the reporting quality of this type of article using the Consolidated Standards of Reporting Trials (CONSORT). We create CONSORT-QA, an evaluation corpus from two studies on abstract reporting quality with CONSORT-abstract standards. We then evaluate the ability of different large generative language models (from the general domain or adapted to the biomedical domain) to correctly assess CONSORT criteria with different known prompting methods, including Chain-of-thought. Our best combination of model and prompting method achieves 85% accuracy. Using Chain-of-thought adds valuable information on the model's reasoning for completing the task.
Explore related subjects
Keep this discovery
Mathieu Laï-king, Patrick Paroubek. 2025-10-05. Evaluation of Clinical Trials Reporting Quality using Large Language Models. https://arxiv.org/abs/2510.04338
Cite the original work for its findings. Save a collection to share your selection of sources.