paper-with-me

홈 › Papers

Reasoning Promotes Robustness in Theory of Mind Tasks

2026-01-23 · Ian B. de Haan, Peter van der Putten, Max van Duijn arxiv

Large language models (LLMs) have recently shown strong performance on Theory of Mind (ToM) tests, prompting debate about the nature and true performance of the underlying capabilities. At the same time, reasoning-oriented LLMs trained via reinforcement learning with verifiable rewards (RLVR) have achieved notable improvements across a range of benchmarks. This paper examines the behavior of such reasoning models in ToM tasks, using novel adaptations of machine psychological experiments and results from established benchmarks. We observe that reasoning models consistently exhibit increased robustness to prompt variations and task perturbations. Our analysis indicates that the observed gains are more plausibly attributed to increased robustness in finding the correct solution, rather than to fundamentally new forms of ToM reasoning. We discuss the implications of this interpretation for evaluating social-cognitive behavior in LLMs.

📄 PDF Abstract BibTeX arXiv:2601.16853

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Understanding Artificial Theory of Mind: Perturbed Tasks and Reasoning in Large Language Models

2026-02-25 · Christian Nickel, Laura Schrewe, Florian Mai, Lucie Flek arxiv

Theory of Mind (ToM) refers to an agent's ability to model the internal states of others. Contributing to the debate whether large language models (LLMs) exhibit genuine ToM capabilities, our study investigates their ToM…

Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

2023-02-16 · Tomer Ullman

Intuitive psychology is a pillar of common-sense reasoning. The replication of this reasoning in machine intelligence is an important stepping-stone on the way to human-like artificial intelligence. Several recent tasks …

Common Sense Reasoning

Theory of Mind with Guilt Aversion Facilitates Cooperative Reinforcement Learning

2020-09-16 · Dung Nguyen, Svetha Venkatesh, Phuoc Nguyen, Truyen Tran

Guilt aversion induces experience of a utility loss in people if they believe they have disappointed others, and this promotes cooperative behaviour in human. In psychological game theory, guilt aversion necessitates mod…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Hypothesis-Driven Theory-of-Mind Reasoning for Large Language Models

2025-02-17 · Hyunwoo Kim, Melanie Sclar, Tan Zhi-Xuan, Lance Ying 외

Existing LLM reasoning methods have shown impressive capabilities across various tasks, such as solving math and coding problems. However, applying these methods to scenarios without ground-truth answers or rule-based ve…

Math

Incentivizing Cardiologist-Like Reasoning in MLLMs for Interpretable Echocardiographic Diagnosis

2026-01-13 · Yi Qin, Lehan Wang, Chenxu Zhao, Alex P. W. Lee 외 arxiv

Echocardiographic diagnosis is vital for cardiac screening yet remains challenging. Existing echocardiography foundation models do not effectively capture the relationships between quantitative measurements and clinical …

Reinforcement Learning