paper-with-me

홈 › Papers

Large language models show fragile cognitive reasoning about human emotions

2025-08-07 · Sree Bhattacharyya, Evgenii Kuriabov, Lucas Craig, Tharun Dilliraj, Reginald B. Adams,, Jia Li, James Z. Wang arxiv

Affective computing seeks to support the holistic development of artificial intelligence by enabling machines to engage with human emotion. Recent foundation models, particularly large language models (LLMs), have been trained and evaluated on emotion-related tasks, typically using supervised learning with discrete emotion labels. Such evaluations largely focus on surface phenomena, such as recognizing expressed or evoked emotions, leaving open whether these systems reason about emotion in cognitively meaningful ways. Here we ask whether LLMs can reason about emotions through underlying cognitive dimensions rather than labels alone. Drawing on cognitive appraisal theory, we introduce CoRE, a large-scale benchmark designed to probe the implicit cognitive structures LLMs use when interpreting emotionally charged situations. We assess alignment with human appraisal patterns, internal consistency, cross-model generalization, and robustness to contextual variation. We find that LLMs capture systematic relations between cognitive appraisals and emotions but show misalignment with human judgments and instability across contexts.

📄 PDF Abstract BibTeX arXiv:2508.05880

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Latent Reward Steering: An Adaptive Inference-Time Framework that Implicitly Promotes Cognitive Behaviors in Reasoning LLMs

2026-05-30 · Jiakang Li, Guanyu Zhu, Can Jin, Chenxi Huang 외 arxiv

Strong reasoning depends not only on model knowledge but also on how effectively cognitive behaviors are deployed during generation. Existing methods often rely on explicit behavior-level control, making them insufficien…

MME-CC: A Challenging Multi-Modal Evaluation Benchmark of Cognitive Capacity

2025-11-05 · Kaiyuan Zhang, Chenghao Yang, Zhoufutu Wen, Sihang Yuan 외 arxiv

As reasoning models scale rapidly, the essential role of multimodality in human cognition has come into sharp relief, driving a growing need to probe vision-centric cognitive behaviors. Yet, existing multimodal benchmark…

GAMBIT: A Gamified Jailbreak Framework for Multimodal Large Language Models

2026-01-06 · Xiangdong Hu, Yangyang Jiang, Qin Hu, Xiaojun Jia arxiv

Multimodal Large Language Models (MLLMs) have become widely deployed, yet their safety alignment remains fragile under adversarial inputs. Previous work has shown that increasing inference steps can disrupt safety mechan…

Cognitive Alpha Mining via LLM-Driven Code-Based Evolution

2025-11-24 · Fengyuan Liu, Yi Huang, Sichun Luo, Yuqi Wang 외 arxiv

Discovering effective predictive signals, or "alphas," from financial data with high dimensionality and extremely low signal-to-noise ratio remains a difficult open problem. Despite progress in deep learning, genetic pro…

Stable Reasoning, Unstable Responses: Mitigating LLM Deception via Stability Asymmetry

2026-03-27 · Guoxi Zhang, Jiawei Chen, Tianzhuo Yang, Lang Qin 외 arxiv

As Large Language Models (LLMs) expand in capability and application scope, their trustworthiness becomes critical. A vital risk is intrinsic deception, wherein models strategically mislead users to achieve their own obj…

Reinforcement Learning