paper-with-me

홈 › Papers

ToMATO: Verbalizing the Mental States of Role-Playing LLMs for Benchmarking Theory of Mind

2025-01-15 · Kazutoshi Shinoda, Nobukatsu Hojo, Kyosuke Nishida, Saki Mizuno, Keita Suzuki, Ryo Masumura, Hiroaki Sugiyama, Kuniko Saito

Existing Theory of Mind (ToM) benchmarks diverge from real-world scenarios in three aspects: 1) they assess a limited range of mental states such as beliefs, 2) false beliefs are not comprehensively explored, and 3) the diverse personality traits of characters are overlooked. To address these challenges, we introduce ToMATO, a new ToM benchmark formulated as multiple-choice QA over conversations. ToMATO is generated via LLM-LLM conversations featuring information asymmetry. By employing a prompting method that requires role-playing LLMs to verbalize their thoughts before each utterance, we capture both first- and second-order mental states across five categories: belief, intention, desire, emotion, and knowledge. These verbalized thoughts serve as answers to questions designed to assess the mental states of characters within conversations. Furthermore, the information asymmetry introduced by hiding thoughts from others induces the generation of false beliefs about various mental states. Assigning distinct personality traits to LLMs further diversifies both utterances and thoughts. ToMATO consists of 5.4k questions, 753 conversations, and 15 personality trait patterns. Our analysis shows that this dataset construction approach frequently generates false beliefs due to the information asymmetry between role-playing LLMs, and effectively reflects diverse personalities. We evaluate nine LLMs on ToMATO and find that even GPT-4o mini lags behind human performance, especially in understanding false beliefs, and lacks robustness to various personality traits.

📄 PDF Abstract BibTeX arXiv:2501.08838

Code (1)

nttmdlab-nlp/ToMATO 공식 구현 pytorch

Tasks

BenchmarkingMultiple-choice

Similar Papers 제목 키워드 기반

On (co-lex) Ordering Automata

2021-06-04 · Giovanna D'Agostino, Nicola Cotumaccio, Alberto Policriti, Nicola Prezza

The states of a deterministic finite automaton A can be identified with collections of words in Pf(L(A)) -- the set of prefixes of words belonging to the regular language accepted by A. But words can be ordered and among…

Reasoning Does Not Necessarily Improve Role-Playing Ability

2025-02-24 · Xiachong Feng, Longxu Dou, Lingpeng Kong

The application of role-playing large language models (LLMs) is rapidly expanding in both academic and commercial domains, driving an increasing demand for high-precision role-playing models. Simultaneously, the rapid ad…

Bounded Rationality in Las Vegas: Probabilistic Finite Automata PlayMulti-Armed Bandits

2020-06-30 · Xinming Liu, Joseph Y. Halpern

While traditional economics assumes that humans are fully rational agents who always maximize their expected utility, in practice, we constantly observe apparently irrational behavior. One explanation is that people have…

Emotional RAG: Enhancing Role-Playing Agents through Emotional Retrieval

2024-10-30 · Le Huang, Hengzhi Lan, Zijun Sun, Chuan Shi 외

As LLMs exhibit a high degree of human-like capability, increasing attention has been paid to role-playing research areas in which responses generated by LLMs are expected to mimic human replies. This has promoted the ex…

RAGResponse GenerationRetrievalRetrieval-augmented Generation+2

Quantum Measurement, Entanglement and the Warping Mechanism of Human Perception

2025-05-01 · Diederik Aerts, Jonito Aerts Arguëlles, Sandro Sozzo

We prove that the quantum measurement process contains the same warping mechanism that occurs in categorical perception, a phenomenon ubiquitous in human perception. This warping causes stimuli belonging to the same cate…