paper-with-me

Papers

TurnaboutLLM: A Deductive Reasoning Benchmark from Detective Games

2025-05-21 · Yuan Yuan, Muyu He, Muhammad Adil Shahid, Jiani Huang, Ziyang Li, Li Zhang

This paper introduces TurnaboutLLM, a novel framework and dataset for evaluating the deductive reasoning abilities of Large Language Models (LLMs) by leveraging the interactive gameplay of detective games Ace Attorney and Danganronpa. The framework tasks LLMs with identifying contradictions between testimonies and evidences within long narrative contexts, a challenging task due to the large answer space and diverse reasoning types presented by its questions. We evaluate twelve state-of-the-art LLMs on the dataset, hinting at limitations of popular strategies for enhancing deductive reasoning such as extensive thinking and Chain-of-Thought prompting. The results also suggest varying effects of context size, the number of reasoning step and answer space size on model performance. Overall, TurnaboutLLM presents a substantial challenge for LLMs' deductive reasoning abilities in complex, narrative-rich environments.

📄 PDF Abstract BibTeX arXiv:2505.15712

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DetectiveQA: Evaluating Long-Context Reasoning on Detective Novels

2024-09-04 · Zhe Xu, Jiasheng Ye, Xiangyang Liu, Tianxiang Sun 외

With the rapid advancement of Large Language Models (LLMs), long-context information understanding and processing have become a hot topic in academia and industry. However, benchmarks for evaluating the ability of LLMs t…

Deciphering Digital Detectives: Understanding LLM Behaviors and Capabilities in Multi-Agent Mystery Games

2023-12-01 · Dekun Wu, Haochen Shi, Zhiyuan Sun, Bang Liu

In this study, we explore the application of Large Language Models (LLMs) in \textit{Jubensha}, a Chinese detective role-playing game and a novel area in Artificial Intelligence (AI) driven gaming. We introduce the first…

AI AgentIn-Context LearningLanguage ModelingLanguage Modelling+2

How Clued up are LLMs? Evaluating Multi-Step Deductive Reasoning in a Text-Based Game Environment

2026-03-17 · Rebecca Ansell, Autumn Toney-Wails arxiv

Deducing whodunit proves challenging for LLM agents. In this paper, we implement a text-based multi-agent version of the classic board game Clue as a rule-based testbed for evaluating multi-step deductive reasoning, with…

Piecing Together Clues: A Benchmark for Evaluating the Detective Skills of Large Language Models

2023-07-11 · Zhouhong Gu, Lin Zhang, Jiangjie Chen, Haoning Ye 외

Detectives frequently engage in information detection and reasoning simultaneously when making decisions across various cases, especially when confronted with a vast amount of information. With the rapid development of l…

Common Sense ReasoningDecision MakingPrompt EngineeringReading Comprehension

Improvisational Games as a Benchmark for Social Intelligence of AI Agents: The Case of Connections

2026-03-31 · Gaurav Rajesh Parikh, Angikar Ghosal arxiv

We formally introduce a improvisational wordplay game called Connections to explore reasoning capabilities of AI agents. Playing Connections combines skills in knowledge retrieval, summarization and awareness of cognitiv…