paper-with-me

Papers

Evaluating and Modeling Social Intelligence: A Comparative Study of Human and AI Capabilities

2024-05-20 · Junqi Wang, Chunhui Zhang, Jiapeng Li, Yuxi Ma, Lixing Niu, Jiaheng Han, Yujia Peng, Yixin Zhu, Lifeng Fan

Facing the current debate on whether Large Language Models (LLMs) attain near-human intelligence levels (Mitchell & Krakauer, 2023; Bubeck et al., 2023; Kosinski, 2023; Shiffrin & Mitchell, 2023; Ullman, 2023), the current study introduces a benchmark for evaluating social intelligence, one of the most distinctive aspects of human cognition. We developed a comprehensive theoretical framework for social dynamics and introduced two evaluation tasks: Inverse Reasoning (IR) and Inverse Inverse Planning (IIP). Our approach also encompassed a computational model based on recursive Bayesian inference, adept at elucidating diverse human behavioral patterns. Extensive experiments and detailed analyses revealed that humans surpassed the latest GPT models in overall performance, zero-shot learning, one-shot generalization, and adaptability to multi-modalities. Notably, GPT models demonstrated social intelligence only at the most basic order (order = 0), in stark contrast to human social intelligence (order >= 2). Further examination indicated a propensity of LLMs to rely on pattern recognition for shortcuts, casting doubt on their possession of authentic human-level social intelligence. Our codes, dataset, appendix and human data are released at https://github.com/bigai-ai/Evaluate-n-Model-Social-Intelligence.

📄 PDF Abstract BibTeX arXiv:2405.11841

Code (1)

bigai-ai/evaluate-n-model-social-intelligence 공식 구현 pytorch

Tasks

Bayesian InferenceZero-Shot Learning

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Attention 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Cosine Annealing Cosine Annealing is a type of learning rate schedule that has the effect of starting with a large learning rate that is relatively rapidly decreased to a minimum value before…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Weight Decay 설명 없음
Linear Warmup With Cosine Annealing Linear Warmup With Cosine Annealing is a learning rate schedule where we increase the learning rate linearly for $n$ updates and then anneal according to a cosine schedule…
Residual Connection 설명 없음

Similar Papers 제목 키워드 기반

The Third Ambition: Artificial Intelligence and the Science of Human Behavior

2026-03-07 · W. Russell Neuman, Chad Coleman arxiv

Contemporary artificial intelligence research has been organized around two dominant ambitions: productivity, which treats AI systems as tools for accelerating work and economic output, and alignment, which focuses on en…

SE-PEF: a Resource for Personalized Expert Finding

2023-09-20 · Pranav Kasela, Gabriella Pasi, Raffaele Perego

The problem of personalization in Information Retrieval has been under study for a long time. A well-known issue related to this task is the lack of publicly available datasets that can support a comparative evaluation o…

Information RetrievalRetrieval

SocialEval: Evaluating Social Intelligence of Large Language Models

2025-06-01 · Jinfeng Zhou, Yuxuan Chen, Yihan Shi, Xuanming Zhang 외

LLMs exhibit promising Social Intelligence (SI) in modeling human behavior, raising the need to evaluate LLMs' SI and their discrepancy with humans. SI equips humans with interpersonal abilities to behave wisely in navig…

Navigate

To Compare, or Not to Compare: On Methodological Practices in Evaluating Social Bias

2026-06-23 · Federico Marcuzzi, Xuefei Ning, Roy Schwartz, Iryna Gurevych arxiv

As Large Language Models are increasingly deployed in critical applications, robustly evaluating their social biases is paramount. However, the current literature suffers from widespread methodological fragmentation, whi…

A Study on the Framework for Evaluating the Ethics and Trustworthiness of Generative AI

2025-08-30 · Cheonsu Jeong, Seunghyun Lee, Seonhee Jeong, Sungsu Kim arxiv

This study provides an in_depth analysis of the ethical and trustworthiness challenges emerging alongside the rapid advancement of generative artificial intelligence (AI) technologies and proposes a comprehensive framewo…