paper-with-me

홈 › Papers

Reasoning Capabilities of Large Language Models. Lessons Learned from General Game Playing

2026-02-22 · Maciej Świechowski, Adam Żychowski, Jacek Mańdziuk arxiv

This paper examines the reasoning capabilities of Large Language Models (LLMs) from a novel perspective, focusing on their ability to operate within formally specified, rule-governed environments. We evaluate four LLMs (Gemini 2.5 Pro and Flash variants, Llama 3.3 70B and GPT-OSS 120B) on a suite of forward-simulation tasks-including next / multistep state formulation, and legal action generation-across a diverse set of reasoning problems illustrated through General Game Playing (GGP) game instances. Beyond reporting instance-level performance, we characterize games based on 40 structural features and analyze correlations between these features and LLM performance. Furthermore, we investigate the effects of various game obfuscations to assess the role of linguistic semantics in game definitions and the impact of potential prior exposure of LLMs to specific games during training. The main results indicate that three of the evaluated models generally perform well across most experimental settings, with performance degradation observed as the evaluation horizon increases (i.e., with a higher number of game steps). Detailed case-based analysis of the LLM performance provides novel insights into common reasoning errors in the considered logic-based problem formulation, including hallucinated rules, redundant state facts, or syntactic errors. Overall, the paper reports clear progress in formal reasoning capabilities of contemporary models.

📄 PDF Abstract BibTeX arXiv:2602.19160

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Toward a Research Agenda in Adversarial Reasoning: Computational Approaches to Anticipating the Opponent's Intent and Actions

2015-12-25 · Alexander Kott, Michael Ownby

This paper defines adversarial reasoning as computational approaches to inferring and anticipating an enemy's perceptions, intents and actions. It argues that adversarial reasoning transcends the boundaries of game theor…

BaRT: A Bayesian Reasoning Tool for Knowledge Based Systems

2013-03-27 · Lashon B. Booker, Naveen Hota, Connie Loggia Ramsey

As the technology for building knowledge based systems has matured, important lessons have been learned about the relationship between the architecture of a system and the nature of the problems it is intended to solve. …

Harnessing the Power of Multiple Minds: Lessons Learned from LLM Routing

2024-05-01 · KV Aditya Srivatsa, Kaushal Kumar Maurya, Ekaterina Kochmar

With the rapid development of LLMs, it is natural to ask how to harness their capabilities efficiently. In this paper, we explore whether it is feasible to direct each input query to a single most suitable LLM. To this e…

Test-time Scaling Techniques in Theoretical Physics -- A Comparison of Methods on the TPBench Dataset

2025-06-25 · Zhiqi Gao, Tianyi Li, Yurii Kvasiuk, Sai Chaitanya Tadepalli 외

Large language models (LLMs) have shown strong capabilities in complex reasoning, and test-time scaling techniques can enhance their performance with comparably low cost. Many of these methods have been developed and eva…

Mathematical Reasoning

Lessons Learned from GPT-SW3: Building the First Large-Scale Generative Language Model for Swedish

2022-06-01 · LREC 2022 6 · Ariel Ekgren, Amaru Cuba Gyllensten, Evangelia Gogoulou, Alice Heiman 외

We present GTP-SW3, a 3.5 billion parameter autoregressive language model, trained on a newly created 100 GB Swedish corpus. This paper provides insights with regards to data collection and training, while highlights the…

Language ModelingLanguage ModellingText Generation