paper-with-me

홈 › Papers

Evaluating Strategic Reasoning in Forecasting Agents

2026-04-28 · Tom Liptay, Dan Schwarz, Rafael Poyiadzi, Jack Wildman, Nikos I. Bosse arxiv

Forecasting benchmarks produce accuracy leaderboards but little insight into why some forecasters are more accurate than others. We introduce Bench to the Future 2 (BTF-2), 1,417 pastcasting questions with a frozen 15M-document research corpus in which agents reproducibly research and forecast offline, producing full reasoning traces. BTF-2 detects accuracy differences of 0.004 Brier score, and can distinguish differential agent strengths in research vs. judgment. We build a forecaster 0.011 Brier more accurate than any single frontier agent, and use it to evaluate agent strategic reasoning without hindsight bias. We find the better forecaster differs primarily in its pre-mortem analysis of its blind spots and consideration of black swans. Expert human forecasters found the dominant strategic reasoning failures of frontier agents are in assessing political and business leaders' incentives, judging their likelihood to follow through on stated plans, and modeling institutional processes.

📄 PDF Abstract BibTeX arXiv:2604.26106

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents

2024-06-07 · Anthony Costarelli, Mat Allen, Roman Hauksson, Grace Sodunke 외

Large language models have demonstrated remarkable few-shot performance on many natural language understanding tasks. Despite several demonstrations of using large language models in complex, strategic scenarios, there l…

Natural Language Understanding

CHBench: A Cognitive Hierarchy Benchmark for Evaluating Strategic Reasoning Capability of LLMs

2025-08-16 · Hongtao Liu, Zhicheng Du, Zihe Wang, Weiran Shen arxiv

Game-playing ability serves as an indicator for evaluating the strategic reasoning capability of large language models (LLMs). While most existing studies rely on utility performance metrics, which are not robust enough …

VS-Bench: Evaluating VLMs for Strategic Reasoning and Decision-Making in Multi-Agent Environments

2025-06-03 · Zelai Xu, Zhexuan Xu, Xiangmin Yi, Huining Yuan 외

Recent advancements in Vision Language Models (VLMs) have expanded their capabilities to interactive agent tasks, yet existing benchmarks remain limited to single-agent or text-only environments. In contrast, real-world …

Decision Making

The Influence of Human-inspired Agentic Sophistication in LLM-driven Strategic Reasoners

2025-05-14 · Vince Trencsenyi, Agnieszka Mensfelt, Kostas Stathis

The rapid rise of large language models (LLMs) has shifted artificial intelligence (AI) research toward agentic systems, motivating the use of weaker and more flexible notions of agency. However, this shift raises key qu…

K-Level Reasoning: Establishing Higher Order Beliefs in Large Language Models for Strategic Reasoning

2024-02-02 · Yadong Zhang, Shaoguang Mao, Tao Ge, Xun Wang 외

Strategic reasoning is a complex yet essential capability for intelligent agents. It requires Large Language Model (LLM) agents to adapt their strategies dynamically in multi-agent environments. Unlike static reasoning t…

Decision MakingLanguage ModellingLarge Language Model