paper-with-me

Papers

NegotiationToM: A Benchmark for Stress-testing Machine Theory of Mind on Negotiation Surrounding

2024-04-21 · Chunkit Chan, Cheng Jiayang, Yauwai Yim, Zheye Deng, Wei Fan, Haoran Li, Xin Liu, Hongming Zhang, Weiqi Wang, Yangqiu Song

Large Language Models (LLMs) have sparked substantial interest and debate concerning their potential emergence of Theory of Mind (ToM) ability. Theory of mind evaluations currently focuses on testing models using machine-generated data or game settings prone to shortcuts and spurious correlations, which lacks evaluation of machine ToM ability in real-world human interaction scenarios. This poses a pressing demand to develop new real-world scenario benchmarks. We introduce NegotiationToM, a new benchmark designed to stress-test machine ToM in real-world negotiation surrounding covered multi-dimensional mental states (i.e., desires, beliefs, and intentions). Our benchmark builds upon the Belief-Desire-Intention (BDI) agent modeling theory and conducts the necessary empirical experiments to evaluate large language models. Our findings demonstrate that NegotiationToM is challenging for state-of-the-art LLMs, as they consistently perform significantly worse than humans, even when employing the chain-of-thought (CoT) method.

📄 PDF Abstract BibTeX arXiv:2404.13627

Code (1)

HKUST-KnowComp/NegotiationToM 공식 구현

Similar Papers 제목 키워드 기반

FANToM: A Benchmark for Stress-testing Machine Theory of Mind in Interactions

2023-10-24 · Hyunwoo Kim, Melanie Sclar, Xuhui Zhou, Ronan Le Bras 외

Theory of mind (ToM) evaluations currently focus on testing models using passive narratives that inherently lack interactivity. We introduce FANToM, a new benchmark designed to stress-test ToM within information-asymmetr…

Question Answering

Clever Hans or Neural Theory of Mind? Stress Testing Social Reasoning in Large Language Models

2023-05-24 · Natalie Shapira, Mosh Levy, Seyed Hossein Alavi, Xuhui Zhou 외

The escalating debate on AI's capabilities warrants developing reliable metrics to assess machine "intelligence". Recently, many anecdotal examples were used to suggest that newer large language models (LLMs) like ChatGP…

The Adaptive Stress Testing Formulation

2020-04-08 · Mark Koren, Anthony Corso, Mykel J. Kochenderfer

Validation is a key challenge in the search for safe autonomy. Simulations are often either too simple to provide robust validation, or too complex to tractably compute. Therefore, approximate validation methods are need…

Machine Learning Based Stress Testing Framework for Indian Financial Market Portfolios

2025-07-02 · Vidya Sagar G, Shifat Ali, Siddhartha P. Chakrabarty arxiv

This paper presents a machine learning driven framework for sectoral stress testing in the Indian financial market, focusing on financial services, information technology, energy, consumer goods, and pharmaceuticals. Ini…

Dimensionality Reduction

Stress-testing medical large language models reveals latent safety pathology beyond benchmark accuracy

2026-06-06 · Yuan Shen, Xiaojun Wu, Linghua Yu arxiv

Large language models (LLMs) are entering clinical practice based on benchmark accuracy that may fail to detect safety-relevant failure modes. Here we present AI-MASLD, a stress-audit framework that adapts the logic of m…

Information Extraction