paper-with-me

Papers

Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human States

2025-05-23 · Yang Xiao, Jiashuo Wang, Qiancheng Xu, Changhe Song, Chunpu Xu, Yi Cheng, Wenjie Li, PengFei Liu

As Large Language Models (LLMs) increasingly participate in human-AI interactions, evaluating their Theory of Mind (ToM) capabilities - particularly their ability to track dynamic mental states - becomes crucial. While existing benchmarks assess basic ToM abilities, they predominantly focus on static snapshots of mental states, overlooking the temporal evolution that characterizes real-world social interactions. We present \textsc{DynToM}, a novel benchmark specifically designed to evaluate LLMs' ability to understand and track the temporal progression of mental states across interconnected scenarios. Through a systematic four-step framework, we generate 1,100 social contexts encompassing 5,500 scenarios and 78,100 questions, each validated for realism and quality. Our comprehensive evaluation of ten state-of-the-art LLMs reveals that their average performance underperforms humans by 44.7\%, with performance degrading significantly when tracking and reasoning about the shift of mental states. This performance gap highlights fundamental limitations in current LLMs' ability to model the dynamic nature of human mental states.

📄 PDF Abstract BibTeX arXiv:2505.17663

Code (1)

GAIR-NLP/DynToM 공식 구현

Tasks

Theory of Mind Modeling

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

InMind: Evaluating LLMs in Capturing and Applying Individual Human Reasoning Styles

2025-08-22 · Zizhen Li, Chuanhao Li, Yibin Wang, Qi Chen 외 arxiv

LLMs have shown strong performance on human-centric reasoning tasks. While previous evaluations have explored whether LLMs can infer intentions or detect deception, they often overlook the individualized reasoning styles…

Can Large Language Models Adapt to Other Agents In-Context?

2024-12-27 · Matthew Riemer, Zahra Ashktorab, Djallel Bouneffouf, Payel Das 외

As the research community aims to build better AI assistants that are more dynamic and personalized to the diversity of humans that they interact with, there is increased interest in evaluating the theory of mind capabil…

Inductive Bias

Reasoning Promotes Robustness in Theory of Mind Tasks

2026-01-23 · Ian B. de Haan, Peter van der Putten, Max van Duijn arxiv

Large language models (LLMs) have recently shown strong performance on Theory of Mind (ToM) tests, prompting debate about the nature and true performance of the underlying capabilities. At the same time, reasoning-orient…

Reinforcement Learning

UrbanMind: Urban Dynamics Prediction with Multifaceted Spatial-Temporal Large Language Models

2025-05-16 · Yuhang Liu, Yingxue Zhang, Xin Zhang, Ling Tian 외

Understanding and predicting urban dynamics is crucial for managing transportation systems, optimizing urban planning, and enhancing public services. While neural network-based approaches have achieved success, they ofte…

Test-time Adaptation

Evaluating Theory of Mind in Question Answering

2018-08-28 · EMNLP 2018 10 · Aida Nematzadeh, Kaylee Burns, Erin Grant, Alison Gopnik 외

We propose a new dataset for evaluating question answering models with respect to their capacity to reason about beliefs. Our tasks are inspired by theory-of-mind experiments that examine whether children are able to rea…

Question Answering