paper-with-me

Papers

TTT-Bench: A Benchmark for Evaluating Reasoning Ability with Simple and Novel Tic-Tac-Toe-style Games

2025-06-11 · Prakamya Mishra, Jiang Liu, Jialian Wu, Xiaodong Yu, Zicheng Liu, Emad Barsoum

Large reasoning models (LRMs) have demonstrated impressive reasoning capabilities across a broad range of tasks including Olympiad-level mathematical problems, indicating evidence of their complex reasoning abilities. While many reasoning benchmarks focus on the STEM domain, the ability of LRMs to reason correctly in broader task domains remains underexplored. In this work, we introduce \textbf{TTT-Bench}, a new benchmark that is designed to evaluate basic strategic, spatial, and logical reasoning abilities in LRMs through a suite of four two-player Tic-Tac-Toe-style games that humans can effortlessly solve from a young age. We propose a simple yet scalable programmatic approach for generating verifiable two-player game problems for TTT-Bench. Although these games are trivial for humans, they require reasoning about the intentions of the opponent, as well as the game board's spatial configurations, to ensure a win. We evaluate a diverse set of state-of-the-art LRMs, and \textbf{discover that the models that excel at hard math problems frequently fail at these simple reasoning games}. Further testing reveals that our evaluated reasoning models score on average $\downarrow$ 41\% \& $\downarrow$ 5\% lower on TTT-Bench compared to MATH 500 \& AIME 2024 respectively, with larger models achieving higher performance using shorter reasoning traces, where most of the models struggle on long-term strategic reasoning situations on simple and new TTT-Bench tasks.

📄 PDF Abstract BibTeX arXiv:2506.10209

Code (0)

등록된 구현이 없습니다.

Tasks

Logical ReasoningMath

Methods 이 논문이 사용한 방법론

Focus 설명 없음
SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

S1-Bench: A Simple Benchmark for Evaluating System 1 Thinking Capability of Large Reasoning Models

2025-04-14 · Wenyuan Zhang, Shuaiyi Nie, Xinghua Zhang, Zefeng Zhang 외

We introduce S1-Bench, a novel benchmark designed to evaluate the performance of Large Reasoning Models (LRMs) on simple tasks that favor intuitive system 1 thinking rather than deliberative system 2 reasoning. While LRM…

Natural Questions

L0-Reasoning Bench: Evaluating Procedural Correctness in Language Models via Simple Program Execution

2025-03-28 · Simeng Sun, Cheng-Ping Hsieh, Faisal Ladhak, Erik Arakelyan 외

Complex reasoning tasks often rely on the ability to consistently and accurately apply simple rules across incremental steps, a foundational capability which we term "level-0" reasoning. To systematically evaluate this c…

KoSimpleQA: A Korean Factuality Benchmark with an Analysis of Reasoning LLMs

2025-10-21 · Donghyeon Ko, Yeguk Jin, Kyubyung Chae, Byungwook Lee 외 arxiv

We present $\textbf{Korean SimpleQA (KoSimpleQA)}$, a benchmark for evaluating factuality in large language models (LLMs) with a focus on Korean cultural knowledge. KoSimpleQA is designed to be challenging yet easy to gr…

LLMs for Relational Reasoning: How Far are We?

2024-01-17 · Zhiming Li, Yushi Cao, Xiufeng Xu, Junzhe Jiang 외

Large language models (LLMs) have revolutionized many areas (e.g. natural language processing, software engineering, etc.) by achieving state-of-the-art performance on extensive downstream tasks. Aiming to achieve robust…

Common Sense ReasoningDecision MakingInductive logic programmingProgram induction+2

MMR: Evaluating Reading Ability of Large Multimodal Models

2024-08-26 · Jian Chen, Ruiyi Zhang, Yufan Zhou, Ryan Rossi 외

Large multimodal models (LMMs) have demonstrated impressive capabilities in understanding various types of image, including text-rich images. Most existing text-rich image benchmarks are simple extraction-based question …

Font RecognitionMMR totalOptical Character Recognition (OCR)Question Answering+2