paper-with-me

홈 › Papers

Evaluating LLMs on Real-World Forecasting Against Expert Forecasters

2025-07-06 · Janna Lu arxiv

Large language models (LLMs) have demonstrated remarkable capabilities across diverse tasks, but their ability to forecast future events remains understudied. A year ago, large language models struggle to come close to the accuracy of a human crowd. I evaluate state-of-the-art LLMs on 464 forecasting questions from Metaculus, comparing their performance against top forecasters. Frontier models achieve Brier scores that ostensibly surpass the human crowd but still significantly underperform a group of experts.

📄 PDF Abstract BibTeX arXiv:2507.04562

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Macroeconomic Forecasting with Large Language Models

2024-07-01 · Andrea Carriero, Davide Pettenuzzo, Shubhranshu Shekhar

This paper presents a comparative analysis evaluating the accuracy of Large Language Models (LLMs) against traditional macro time series forecasting approaches. In recent times, LLMs have surged in popularity for forecas…

Time SeriesTime Series Forecasting

WorldReasoner: Evaluating Whether Language Model Agents Forecast Events with Valid Reasoning

2026-06-10 · Yizhou Chi, Eric Chamoun, Zifeng Ding, Andreas Vlachos arxiv

Forecasting real-world events requires language-model agents to reason under uncertainty from incomplete, time-bounded information. Yet evaluating whether agents genuinely forecast requires more than final-answer accurac…

LEAF: A Living Benchmark for Event-Augmented Forecasting

2026-05-09 · Mingtian Tan, Mihir Parmar, Palash Goyal, Chun-Liang Li 외 arxiv

Large Language Models (LLMs) are increasingly applied to forecasting. To evaluate this capability while mitigating pre-training data contamination, several living benchmarks have been proposed. However, existing benchmar…

Time Series Forecasting

MultiCast: Zero-Shot Multivariate Time Series Forecasting Using LLMs

2024-05-23 · Georgios Chatzigeorgakidis, Konstantinos Lentzos, Dimitrios Skoutas

Predicting future values in multivariate time series is vital across various domains. This work explores the use of large language models (LLMs) for this task. However, LLMs typically handle one-dimensional data. We intr…

Multivariate Time Series ForecastingQuantizationTime SeriesTime Series Forecasting

Automating Forecasting Question Generation and Resolution for AI Evaluation

2026-01-30 · Nikos I. Bosse, Peter Mühlbacher, Jack Wildman, Lawrence Phillips 외 arxiv

Forecasting future events is highly valuable in decision-making and is a robust measure of general intelligence. As forecasting is probabilistic, developing and evaluating AI forecasters requires generating large numbers…

Question Generation