paper-with-me

Papers

Towards Objectively Benchmarking Social Intelligence for Language Agents at Action Level

2024-04-08 · Chenxu Wang, Bin Dai, Huaping Liu, Baoyuan Wang

Prominent large language models have exhibited human-level performance in many domains, even enabling the derived agents to simulate human and social interactions. While practical works have substantiated the practicability of grounding language agents in sandbox simulation or embodied simulators, current social intelligence benchmarks either stay at the language level or use subjective metrics. In pursuit of a more realistic and objective evaluation, we introduce the Social Tasks in Sandbox Simulation (STSS) benchmark, which assesses language agents \textbf{objectively} at the \textbf{action level} by scrutinizing the goal achievements within the multi-agent simulation. Additionally, we sample conversation scenarios to build a language-level benchmark to provide an economically prudent preliminary evaluation and align with prevailing benchmarks. To gauge the significance of agent architecture, we implement a target-driven planning (TDP) module as an adjunct to the existing agent. Our evaluative findings highlight that the STSS benchmark is challenging for state-of-the-art language agents. Furthermore, it effectively discriminates between distinct language agents, suggesting its usefulness as a benchmark for evaluating both language models and agent architectures.

📄 PDF Abstract BibTeX arXiv:2404.05337

Code (1)

wcx21/Social-Tasks-in-Sandbox-Simulation 공식 구현

Tasks

Benchmarking

Similar Papers 제목 키워드 기반

AgentSense: Benchmarking Social Intelligence of Language Agents through Interactive Scenarios

2024-10-25 · Xinyi Mou, Jingcong Liang, Jiayu Lin, Xinnong Zhang 외

Large language models (LLMs) are increasingly leveraged to empower autonomous agents to simulate human beings in various fields of behavioral research. However, evaluating their capacity to navigate complex social intera…

BenchmarkingDiversityNavigate

SocialAI: Benchmarking Socio-Cognitive Abilities in Deep Reinforcement Learning Agents

2021-07-02 · Grgur Kovač, Rémy Portelas, Katja Hofmann, Pierre-Yves Oudeyer

Building embodied autonomous agents capable of participating in social interactions with humans is one of the main challenges in AI. Within the Deep Reinforcement Learning (DRL) field, this objective motivated multiple w…

BenchmarkingDeep Reinforcement Learningreinforcement-learningReinforcement Learning+1

Can Agents Read the Room? Benchmarking Visual Social Intelligence in Multimodal Simulation

2026-06-13 · Shijun Wan, Xuehai Wu, Jiwen Zhang, Siyuan Wang 외 arxiv

Social interaction depends on both language and visible social signals, such as facial expressions, posture, gaze, and emotional shifts. Yet existing social-agent benchmarks are largely text-based and rarely test whether…

SI-Bench: Benchmarking Social Intelligence of Large Language Models in Human-to-Human Conversations

2025-10-27 · Shuai Huang, Wenxuan Zhao, Jun Gao arxiv

As large language models (LLMs) develop anthropomorphic abilities, they are increasingly being deployed as autonomous agents to interact with humans. However, evaluating their performance in realistic and complex social …

SOTOPIA-$π$: Interactive Learning of Socially Intelligent Language Agents

2024-03-13 · Ruiyi Wang, Haofei Yu, Wenxin Zhang, Zhengyang Qi 외

Humans learn social skills through both imitation and social interaction. This social learning process is largely understudied by existing research on building language agents. Motivated by this gap, we propose an intera…

Language ModelingLanguage ModellingLarge Language ModelMMLU