paper-with-me

Papers

Does Theory of Mind Improvement Really Benefit Human-AI Interactions? Empirical Findings from Interactive Evaluations

2026-04-28 · Nanxu Gong, Zixin Chen, Haotian Li, Zishu Zhao, Jianxun Lian, Huamin Qu, Yanjie Fu, Xing Xie arxiv

Improving the Theory of Mind (ToM) capability of Large Language Models (LLMs) is crucial for effective social interactions between these AI models and humans. However, the existing benchmarks often measure ToM capability improvement through story-reading, multiple-choice questions from a third-person perspective, while ignoring the first-person, dynamic, and open-ended nature of human-AI (HAI) interactions. To directly examine how ToM improvement techniques benefit HAI interactions, we first proposed the new paradigm of interactive ToM evaluation with both perspective and metric shifts. Next, following the paradigm, we conducted a systematic study of four representative ToM enhancement techniques using both four real-world datasets and a user study, covering both goal-oriented tasks (e.g., coding, math) and experience-oriented tasks (e.g., counseling). Our findings reveal that improvements on static benchmarks do not always translate to better performance in dynamic HAI interactions. This paper offers critical insights into ToM evaluation, showing the necessity of interaction-based assessments in developing next-generation, socially aware LLMs for HAI symbiosis.

📄 PDF Abstract BibTeX arXiv:2605.15205

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MindGames: Targeting Theory of Mind in Large Language Models with Dynamic Epistemic Modal Logic

2023-05-05 · Damien Sileo, Antoine Lernould

Theory of Mind (ToM) is a critical component of intelligence but its assessment remains the subject of heated debates. Prior research applied human ToM assessments to natural language processing models using either human…

Epistemic ReasoningLanguage ModelingLanguage ModellingMultiple-choice

When are ensembles really effective?

2023-05-21 · NeurIPS 2023 11

Ensembling has a long history in statistical data analysis, with many impactful applications. However, in many modern machine learning settings, the benefits of ensembling are less ubiquitous and less obvious. We study, …

A Computable Game-Theoretic Framework for Multi-Agent Theory of Mind

2025-11-27 · Fengming Zhu, Yuxin Pan, Xiaomeng Zhu, Fangzhen Lin arxiv

Originating in psychology, $\textit{Theory of Mind}$ (ToM) has attracted significant attention across multiple research communities, especially logic, economics, and robotics. Most psychological work does not aim at form…

Rethinking ValueDice: Does It Really Improve Performance?

2022-01-17 · ICLR Track Blog 2022 5 · Anonymous

Since the introduction of GAIL, adversarial imitation learning (AIL) methods attract lots of research interests. Among these methods, ValueDice has achieved significant improvements: it beats the classical approach Behav…

Imitation Learning

Rethinking ValueDice: Does It Really Improve Performance?

2022-02-05 · Ziniu Li, Tian Xu, Yang Yu, Zhi-Quan Luo

Since the introduction of GAIL, adversarial imitation learning (AIL) methods attract lots of research interests. Among these methods, ValueDice has achieved significant improvements: it beats the classical approach Behav…

Imitation Learning