arXiv · 2504.20276
Enhancing Systematic Reviews with Large Language Models: Using GPT-4 and Kimi
Abstract
This research delved into GPT-4 and Kimi, two Large Language Models (LLMs), for systematic reviews. We evaluated their performance by comparing LLM-generated codes with human-generated codes from a peer-reviewed systematic review on assessment. Our findings suggested that the performance of LLMs fluctuates by data volume and question complexity for systematic reviews.
Explore related subjects
Keep this discovery
Dandan Chen Kaptur, Yue Huang, Xuejun Ryan Ji, Yanhui Guo, Bradley Kaptur. 2025-04-28. Enhancing Systematic Reviews with Large Language Models: Using GPT-4 and Kimi. https://arxiv.org/abs/2504.20276
Cite the original work for its findings. Save a collection to share your selection of sources.