paper-with-me

Papers

FANToM: A Benchmark for Stress-testing Machine Theory of Mind in Interactions

2023-10-24 · Hyunwoo Kim, Melanie Sclar, Xuhui Zhou, Ronan Le Bras, Gunhee Kim, Yejin Choi, Maarten Sap

Theory of mind (ToM) evaluations currently focus on testing models using passive narratives that inherently lack interactivity. We introduce FANToM, a new benchmark designed to stress-test ToM within information-asymmetric conversational contexts via question answering. Our benchmark draws upon important theoretical requisites from psychology and necessary empirical considerations when evaluating large language models (LLMs). In particular, we formulate multiple types of questions that demand the same underlying reasoning to identify illusory or false sense of ToM capabilities in LLMs. We show that FANToM is challenging for state-of-the-art LLMs, which perform significantly worse than humans even with chain-of-thought reasoning or fine-tuning.

📄 PDF Abstract BibTeX arXiv:2310.15421

Code (0)

등록된 구현이 없습니다.

Tasks

Question Answering

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

NegotiationToM: A Benchmark for Stress-testing Machine Theory of Mind on Negotiation Surrounding

2024-04-21 · Chunkit Chan, Cheng Jiayang, Yauwai Yim, Zheye Deng 외

Large Language Models (LLMs) have sparked substantial interest and debate concerning their potential emergence of Theory of Mind (ToM) ability. Theory of mind evaluations currently focuses on testing models using machine…

OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind

2026-05-19 · Sharmin Sultana Srishty, Kazi Mahathir Rahman, Malaika Parizat Sakkhi, Samia Shahid Prianna 외 arxiv

Large Language Models (LLMs) perform well on many language tasks, but their Theory of Mind (ToM) reasoning is still uneven in complex social settings. Existing benchmarks, including ExploreToM, do not always test the rec…

Reinforcement Learning

Perceptions to Beliefs: Exploring Precursory Inferences for Theory of Mind in Large Language Models

2024-07-08 · Chani Jung, Dongkwan Kim, Jiho Jin, Jiseon Kim 외

While humans naturally develop theory of mind (ToM), the capability to understand other people's mental states and beliefs, state-of-the-art large language models (LLMs) underperform on simple ToM benchmarks. We posit th…

Clever Hans or Neural Theory of Mind? Stress Testing Social Reasoning in Large Language Models

2023-05-24 · Natalie Shapira, Mosh Levy, Seyed Hossein Alavi, Xuhui Zhou 외

The escalating debate on AI's capabilities warrants developing reliable metrics to assess machine "intelligence". Recently, many anecdotal examples were used to suggest that newer large language models (LLMs) like ChatGP…

The Adaptive Stress Testing Formulation

2020-04-08 · Mark Koren, Anthony Corso, Mykel J. Kochenderfer

Validation is a key challenge in the search for safe autonomy. Simulations are often either too simple to provide robust validation, or too complex to tractably compute. Therefore, approximate validation methods are need…