paper-with-me

홈 › Papers

RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity

2025-09-30 · Jisu Shin, Hoyun Song, Juhyun Oh, Changgeon Ko, Eunsu Kim, Chani Jung, Alice Oh arxiv

People often encounter role conflicts -- social dilemmas where the expectations of multiple roles clash and cannot be simultaneously fulfilled. As large language models (LLMs) increasingly navigate these social dynamics, a critical research question emerges. When faced with such dilemmas, do LLMs prioritize dynamic contextual cues or the learned preferences? To address this, we introduce RoleConflictBench, a novel benchmark designed to measure the contextual sensitivity of LLMs in role conflict scenarios. To enable objective evaluation within this subjective domain, we employ situational urgency as a constraint for decision-making. We construct the dataset through a three-stage pipeline that generates over 13,000 realistic scenarios across 65 roles in five social domains by systematically varying the urgency of competing situations. This controlled setup enables us to quantitatively measure contextual sensitivity, determining whether model decisions align with the situational contexts or are overridden by the learned role preferences. Our analysis of 10 LLMs reveals that models substantially deviate from this objective baseline. Instead of responding to dynamic contextual cues, their decisions are predominantly governed by the preferences toward specific social roles.

📄 PDF Abstract BibTeX arXiv:2509.25897

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RoleCDE:Benchmarking and Mitigating Role-Alignment Trade-offs in Role-Playing Agents

2026-06-01 · Huayi Lai, Shichao Song, Simin Niu, Hanyu Wang 외 arxiv

Role-playing agents(RPAs) are widely used to steer large language models(LLMs) toward role-consistent behavior, yet existing benchmarks mainly evaluate surface-level fidelity and offer limited insight into decision makin…

Decision Making

Can Agents Read the Room? Benchmarking Visual Social Intelligence in Multimodal Simulation

2026-06-13 · Shijun Wan, Xuehai Wu, Jiwen Zhang, Siyuan Wang 외 arxiv

Social interaction depends on both language and visible social signals, such as facial expressions, posture, gaze, and emotional shifts. Yet existing social-agent benchmarks are largely text-based and rarely test whether…

OrgAccess: A Benchmark for Role Based Access Control in Organization Scale LLMs

2025-05-25 · Debdeep Sanyal Umakanta Maharana, Yash Sinha, Hong Ming Tan, Shirish Karande 외

Role-based access control (RBAC) and hierarchical structures are foundational to how information flows and decisions are made within virtually all organizations. As the potential of Large Language Models (LLMs) to serve …

ConflictBench: Evaluating Human-AI Conflict via Interactive and Visually Grounded Environments

2026-03-09 · Weixiang Zhao, Haozhen Li, Yanyan Zhao, xuda zhi 외 arxiv

As large language models (LLMs) evolve into autonomous agents capable of acting in open-ended environments, ensuring behavioral alignment with human values becomes a critical safety concern. Existing benchmarks, focused …

Can LLMs Be CEOs? Benchmarking Strategic Resource Reallocation with Multi-Role Agent Simulation

2026-06-16 · Yuyang Dai, Xueqing Peng, Lingfei Qian, Zhuohan Xie arxiv

Evaluating the decision-making capabilities of large language models (LLMs) is a growing research priority, yet existing benchmarks focus on isolated cognitive tasks such as reasoning, knowledge retrieval, and economic r…