arXiv · 2507.04562
Evaluating LLMs on Real-World Forecasting Against Expert Forecasters
Abstract
Large language models (LLMs) have demonstrated remarkable capabilities across diverse tasks, but their ability to forecast future events remains understudied. A year ago, large language models struggle to come close to the accuracy of a human crowd. I evaluate state-of-the-art LLMs on 464 forecasting questions from Metaculus, comparing their performance against top forecasters. Frontier models achieve Brier scores that ostensibly surpass the human crowd but still significantly underperform a group of experts.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Janna Lu. 2025-07-06. Evaluating LLMs on Real-World Forecasting Against Expert Forecasters. https://arxiv.org/abs/2507.04562
Cite the original work for its findings. Save a collection to share your selection of sources.