paper-with-me

Papers

Humans vs Large Language Models: Judgmental Forecasting in an Era of Advanced AI

2023-12-12 · Mahdi Abolghasemi, Odkhishig Ganbold, Kristian Rotaru

This study investigates the forecasting accuracy of human experts versus Large Language Models (LLMs) in the retail sector, particularly during standard and promotional sales periods. Utilizing a controlled experimental setup with 123 human forecasters and five LLMs, including ChatGPT4, ChatGPT3.5, Bard, Bing, and Llama2, we evaluated forecasting precision through Mean Absolute Percentage Error. Our analysis centered on the effect of the following factors on forecasters performance: the supporting statistical model (baseline and advanced), whether the product was on promotion, and the nature of external impact. The findings indicate that LLMs do not consistently outperform humans in forecasting accuracy and that advanced statistical forecasting models do not uniformly enhance the performance of either human forecasters or LLMs. Both human and LLM forecasters exhibited increased forecasting errors, particularly during promotional periods and under the influence of positive external impacts. Our findings call for careful consideration when integrating LLMs into practical forecasting processes.

📄 PDF Abstract BibTeX arXiv:2312.06941

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Argumentatively Coherent Judgmental Forecasting

2025-07-30 · Deniz Gorur, Antonio Rago, Francesca Toni arxiv

Judgmental forecasting employs human opinions to make predictions about future events, rather than exclusively historical data as in quantitative forecasting. When these opinions form an argumentative structure around fo…

Retrieval- and Argumentation-Enhanced Multi-Agent LLMs for Judgmental Forecasting (Extended Version with Supplementary Material)

2025-10-28 · Deniz Gorur, Antonio Rago, Francesca Toni arxiv

Judgmental forecasting is the task of making predictions about future events based on human judgment. This task can be seen as a form of claim verification, where the claim corresponds to a future event and the task is t…

Argument Mining

Scenario Synthesis and Macroeconomic Risk

2025-05-08 · Tobias Adrian, Domenico Giannone, Matteo Luciani, Mike West

We introduce methodology to bridge scenario analysis and model-based risk forecasting, leveraging their respective strengths in policy settings. Our Bayesian framework addresses the fundamental challenge of reconciling j…

AIA Forecaster: Technical Report

2025-11-10 · Rohan Alur, Bradly C. Stadie, Daniel Kang, Ryan Chen 외 arxiv

This technical report describes the AIA Forecaster, a Large Language Model (LLM)-based system for judgmental forecasting using unstructured data. The AIA Forecaster approach combines three core elements: agentic search o…

QuantSightBench: Evaluating LLM Quantitative Forecasting with Prediction Intervals

2026-04-17 · Jeremy Qin, Maksym Andriushchenko arxiv

Forecasting has become a natural benchmark for reasoning under uncertainty. Yet existing evaluations of large language models remain limited to judgmental tasks in simple formats, such as binary or multiple-choice questi…