paper-with-me

Papers

Foresight Optimization for Strategic Reasoning in Large Language Models

2026-04-15 · Jiashuo Wang, Jiawen Duan, Jian Wang, Kaitao Song, Chunpu Xu, Johnny K. W. Ho, Fenggang Yu, Wenjie Li, Johan F. Hoorn arxiv

Reasoning capabilities in large language models (LLMs) have generally advanced significantly. However, it is still challenging for existing reasoning-based LLMs to perform effective decision-making abilities in multi-agent environments, due to the absence of explicit foresight modeling. To this end, strategic reasoning, the most fundamental capability to anticipate the counterpart's behaviors and foresee its possible future actions, has been introduced to alleviate the above issues. Strategic reasoning is fundamental to effective decision-making in multi-agent environments, yet existing reasoning enhancement methods for LLMs do not explicitly capture its foresight nature. In this work, we introduce Foresight Policy Optimization (FoPO) to enhance strategic reasoning in LLMs, which integrates opponent modeling principles into policy optimization, thereby enabling explicit consideration of both self-interest and counterpart influence. Specifically, we construct two curated datasets, namely Cooperative RSA and Competitive Taboo, equipped with well-designed rules and moderate difficulty to facilitate a systematic investigation of FoPO in a self-play framework. Our experiments demonstrate that FoPO significantly enhances strategic reasoning across LLMs of varying sizes and origins. Moreover, models trained with FoPO exhibit strong generalization to out-of-domain strategic scenarios, substantially outperforming standard LLM reasoning optimization baselines.

📄 PDF Abstract BibTeX arXiv:2604.13592

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Top 10 Open Challenges Steering the Future of Diffusion Language Model and Its Variants

2026-01-20 · Yunhe Wang, Kai Han, Huiling Zhen, Yuchuan Tian 외 arxiv

The paradigm of Large Language Models (LLMs) is currently defined by auto-regressive (AR) architectures, which generate text through a sequential ``brick-by-brick'' process. Despite their success, AR models are inherentl…

Text Generation

Decoding the Human Factor: High Fidelity Behavioral Prediction for Strategic Foresight

2026-02-19 · Ben Yellin, Ehud Ezra, Mark Foreman, Shula Grinapol arxiv

Predicting human decision-making in high-stakes environments remains a central challenge for artificial intelligence. While large language models (LLMs) demonstrate strong general reasoning, they often struggle to genera…

Current Agents Fail to Leverage World Model as Tool for Foresight

2026-01-07 · Cheng Qian, Emre Can Acikgoz, Bingxuan Li, Xiusi Chen 외 arxiv

Agents built on vision-language models increasingly face tasks that demand anticipating future states rather than relying on short-horizon reasoning. Generative world models offer a promising remedy: agents could use the…

Visual Question Answering

METRO: Towards Strategy Induction from Expert Dialogue Transcripts for Non-collaborative Dialogues

2026-04-13 · Haofu Yang, Jiaji Liu, Chen Huang, Faguo Wu 외 arxiv

Developing non-collaborative dialogue agents traditionally requires the manual, unscalable codification of expert strategies. We propose \ours, a method that leverages large language models to autonomously induce both st…

Multidimensional Cybersecurity Framework for Strategic Foresight

2022-02-05 · Cyril Onwubiko, Karim Ouazzane

Cybersecurity is now at the forefront of most organisational digital transformative agendas and National economic, social and political programmes. Hence its impact to society can no longer be seen to be one dimensional.…