paper-with-me

Papers

AgentSense: Benchmarking Social Intelligence of Language Agents through Interactive Scenarios

2024-10-25 · Xinyi Mou, Jingcong Liang, Jiayu Lin, Xinnong Zhang, Xiawei Liu, Shiyue Yang, Rong Ye, Lei Chen, Haoyu Kuang, Xuanjing Huang, Zhongyu Wei

Large language models (LLMs) are increasingly leveraged to empower autonomous agents to simulate human beings in various fields of behavioral research. However, evaluating their capacity to navigate complex social interactions remains a challenge. Previous studies face limitations due to insufficient scenario diversity, complexity, and a single-perspective focus. To this end, we introduce AgentSense: Benchmarking Social Intelligence of Language Agents through Interactive Scenarios. Drawing on Dramaturgical Theory, AgentSense employs a bottom-up approach to create 1,225 diverse social scenarios constructed from extensive scripts. We evaluate LLM-driven agents through multi-turn interactions, emphasizing both goal completion and implicit reasoning. We analyze goals using ERG theory and conduct comprehensive experiments. Our findings highlight that LLMs struggle with goals in complex social scenarios, especially high-level growth needs, and even GPT-4o requires improvement in private information reasoning. Code and data are available at \url{https://github.com/ljcleo/agent_sense}.

📄 PDF Abstract BibTeX arXiv:2410.19346

Code (1)

ljcleo/agent_sense 공식 구현

Tasks

BenchmarkingDiversityNavigate

Similar Papers 제목 키워드 기반

AgentSense: LLMs Empower Generalizable and Explainable Web-Based Participatory Urban Sensing

2025-10-22 · Xusen Guo, Mingxing Peng, Xixuan Hao, Xingchen Zou 외 arxiv

Web-based participatory urban sensing has emerged as a vital approach for modern urban management by leveraging mobile individuals as distributed sensors. However, existing urban sensing systems struggle with limited gen…

AgentSense: Virtual Sensor Data Generation Using LLM Agents in Simulated Home Environments

2025-06-13 · Zikang Leng, Megha Thukral, Yaqi Liu, Hrudhai Rajasekhar 외

A major obstacle in developing robust and generalizable smart home-based Human Activity Recognition (HAR) systems is the lack of large-scale, diverse labeled datasets. Variability in home layouts, sensor configurations, …

Activity RecognitionHuman Activity Recognition

Towards Objectively Benchmarking Social Intelligence for Language Agents at Action Level

2024-04-08 · Chenxu Wang, Bin Dai, Huaping Liu, Baoyuan Wang

Prominent large language models have exhibited human-level performance in many domains, even enabling the derived agents to simulate human and social interactions. While practical works have substantiated the practicabil…

Benchmarking

SocialAI: Benchmarking Socio-Cognitive Abilities in Deep Reinforcement Learning Agents

2021-07-02 · Grgur Kovač, Rémy Portelas, Katja Hofmann, Pierre-Yves Oudeyer

Building embodied autonomous agents capable of participating in social interactions with humans is one of the main challenges in AI. Within the Deep Reinforcement Learning (DRL) field, this objective motivated multiple w…

BenchmarkingDeep Reinforcement Learningreinforcement-learningReinforcement Learning+1

Can Agents Read the Room? Benchmarking Visual Social Intelligence in Multimodal Simulation

2026-06-13 · Shijun Wan, Xuehai Wu, Jiwen Zhang, Siyuan Wang 외 arxiv

Social interaction depends on both language and visible social signals, such as facial expressions, posture, gaze, and emotional shifts. Yet existing social-agent benchmarks are largely text-based and rarely test whether…