paper-with-me

홈 › Papers

LEAF: A Living Benchmark for Event-Augmented Forecasting

2026-05-09 · Mingtian Tan, Mihir Parmar, Palash Goyal, Chun-Liang Li, Nanyun Peng, Thomas Hartvigsen, Jinsung Yoon, Tomas Pfister arxiv

Large Language Models (LLMs) are increasingly applied to forecasting. To evaluate this capability while mitigating pre-training data contamination, several living benchmarks have been proposed. However, existing benchmarks either lack the multidimensional events essential for accurate forecasting due to data scarcity, or focus on relatively closed environments. To assess the predictive capabilities of LLMs in complex, real-world scenarios, we propose LEAF, the first living benchmark for event-augmented forecasting tasks, including future event probabilities, trend and time series forecasting. LEAF utilizes a recursive retrieval agent system paired with dual-agent cross-validation to provide comprehensive and relevant auxiliary text for forecasting. Evaluating state-of-the-art proprietary and open-weight LLMs, we find that these models can leverage signals extracted from complex events to enhance predictive performance. In the stock domain, we find that LLMs achieve better performance on equities they confidently identify as more predictable. Furthermore, the events demonstrate a strong correlation with the target equities. To this end, LEAF provides a necessary, dynamically updating testbed to continuously track and drive progress in event-driven forecasting tasks.

📄 PDF Abstract BibTeX arXiv:2605.16358

Code (0)

등록된 구현이 없습니다.

Tasks

Time Series Forecasting

Similar Papers 제목 키워드 기반

Bench to the Future: A Pastcasting Benchmark for Forecasting Agents

2025-06-11 · FutureSearch, :, Jack Wildman, Nikos I. Bosse 외

Forecasting is a challenging task that offers a clearly measurable way to study AI systems. Forecasting requires a large amount of research on the internet, and evaluations require time for events to happen, making the d…

Benchmarking

SURGE: An Event-Centric Social Media Sentiment Time Series Benchmark with Interaction Structure

2026-05-20 · Chen Su, Pengsen Cheng, Yuanhe Tian, Yan Song arxiv

Public events on social media generate large volumes of discussion whose collective dynamics carry direct value for opinion forecasting and crisis response. Capturing how these dynamics evolve across an event's lifecycle…

A Comprehensive Evaluation of Large Language Models on Temporal Event Forecasting

2024-07-16 · He Chang, Chenchen Ye, Zhulin Tao, Jie Wu 외

Recently, Large Language Models (LLMs) have demonstrated great potential in various data mining tasks, such as knowledge question answering, mathematical reasoning, and commonsense reasoning. However, the reasoning capab…

Mathematical ReasoningQuestion AnsweringRAGRetrieval+1

Novel Regression and Least Square Support Vector Machine Learning Technique for Air Pollution Forecasting

2023-06-11 · Dhanalakshmi M, Radha V

Air pollution is the origination of particulate matter, chemicals, or biological substances that brings pain to either humans or other living creatures or instigates discomfort to the natural habitat and the airspace. He…

regression

PROPHET: An Inferable Future Forecasting Benchmark with Causal Intervened Likelihood Estimation

2025-04-02 · Zhengwei Tao, Zhi Jin, Bincheng Li, Xiaoying Bai 외

Predicting future events stands as one of the ultimate aspirations of artificial intelligence. Recent advances in large language model (LLM)-based systems have shown remarkable potential in forecasting future events, the…

ArticlesCausal InferenceLarge Language ModelPrediction+3