paper-with-me

홈 › Papers

Views Are My Own, but Also Yours: Benchmarking Theory of Mind Using Common Ground

2024-03-04 · Adil Soubki, John Murzaku, Arash Yousefi Jordehi, Peter Zeng, Magdalena Markowska, Seyed Abolghasem Mirroshandel, Owen Rambow

Evaluating the theory of mind (ToM) capabilities of language models (LMs) has recently received a great deal of attention. However, many existing benchmarks rely on synthetic data, which risks misaligning the resulting experiments with human behavior. We introduce the first ToM dataset based on naturally occurring spoken dialogs, Common-ToM, and show that LMs struggle to demonstrate ToM. We then show that integrating a simple, explicit representation of beliefs improves LM performance on Common-ToM.

📄 PDF Abstract BibTeX arXiv:2403.02451

Code (0)

등록된 구현이 없습니다.

Tasks

Benchmarking

Similar Papers 제목 키워드 기반

A formal definition and meta-model for a machine theory of mind

2026-06-02 · Fabio Cuzzolin arxiv

This paper proposes, for the first time, a rigorous formal definition of the concept of Machine Theory of Mind, based on principles supported by evidence from cognitive psychology, neuroscience and artificial intelligenc…

Mind Your Theory: Theory of Mind Goes Deeper Than Reasoning

2024-12-18 · Eitan Wagner, Nitay Alon, Joseph M. Barnby, Omri Abend

Theory of Mind (ToM) capabilities in LLMs have recently become a central object of investigation. Cognitive science distinguishes between two steps required for ToM tasks: 1) determine whether to invoke ToM, which includ…

BenchmarkingPosition

Mind the Motions: Benchmarking Theory-of-Mind in Everyday Body Language

2025-11-19 · Seungbeen Lee, Jinhong Jeong, Donghyun Kim, Yejin Son 외 arxiv

Our ability to interpret others' mental states through nonverbal cues (NVCs) is fundamental to our survival and social cohesion. While existing Theory of Mind (ToM) benchmarks have primarily focused on false-belief tasks…

Minds, Brains, AI

2024-04-21 · Jay Seitz

In the last year or so and going back many decades there has been extensive claims by major computational scientists, engineers, and others that AGI, artificial general intelligence, is five or ten years away, but withou…

Self-Driving Cars

GroupToM-Bench: Benchmarking Group Theory of Mind and Nonlinear Social Emergence in MLLMs

2026-06-02 · Weidong Tang, Jierui Li, Yueling Hou, Zihan Mei 외 arxiv

True general intelligence requires not only a model of the physical world but also a social world model: the capacity to infer how individual mental states interact and crystallize into group-level outcomes. Despite nota…