paper-with-me

홈 › Papers

Role-Conditioned Refusals: Evaluating Access Control Reasoning in Large Language Models

2025-10-09 · Đorđe Klisura, Joseph Khoury, Ashish Kundu, Ram Krishnan, Anthony Rios arxiv

Access control is a cornerstone of secure computing, yet large language models often blur role boundaries by producing unrestricted responses. We study role-conditioned refusals, focusing on the LLM's ability to adhere to access control policies by answering when authorized and refusing when not. To evaluate this behavior, we created a novel dataset that extends the Spider and BIRD text-to-SQL datasets, both of which have been modified with realistic PostgreSQL role-based policies at the table and column levels. We compare three designs: (i) zero or few-shot prompting, (ii) a two-step generator-verifier pipeline that checks SQL against policy, and (iii) LoRA fine-tuned models that learn permission awareness directly. Across multiple model families, explicit verification (the two-step framework) improves refusal precision and lowers false permits. At the same time, fine-tuning achieves a stronger balance between safety and utility (i.e., when considering execution accuracy). Longer and more complex policies consistently reduce the reliability of all systems. We release RBAC-augmented datasets and code.

📄 PDF Abstract BibTeX arXiv:2510.07642

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Jinx: Unlimited LLMs for Probing Alignment Failures

2025-08-11 · Jiahao Zhao, Liwei Dong arxiv

Unlimited, or so-called helpful-only language models are trained without safety alignment constraints and never refuse user queries. They are widely used by leading AI companies as internal tools for red teaming and alig…

Instruction FollowingRed Teaming

Equal Access, Unequal Interaction: A Counterfactual Audit of LLM Fairness

2026-02-03 · Alireza Amiri-Margavi, Arshia Gharagozlou, Amin Gholami Davodi, Seyed Pouyan Mousavi Davoudi 외 arxiv

Prior work on fairness in large language models (LLMs) has primarily focused on access-level behaviors such as refusals and safety filtering. However, equitable access does not ensure equitable interaction quality once a…

Role-Aware Language Models for Secure and Contextualized Access Control in Organizations

2025-07-31 · Saeed Almheiri, Yerulan Kongrat, Adrian Santosh, Ruslan Tasmukhanov 외 arxiv

As large language models (LLMs) are increasingly deployed in enterprise settings, controlling model behavior based on user roles becomes an essential requirement. Existing safety methods typically assume uniform access a…

Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models

2024-09-01 · Bang An, Sicheng Zhu, Ruiyi Zhang, Michael-Andrei Panaitescu-Liess 외

Safety-aligned large language models (LLMs) sometimes falsely refuse pseudo-harmful prompts, like "how to kill a mosquito," which are actually harmless. Frequent false refusals not only frustrate users but also provoke a…

ValueGround: Evaluating Culture-Conditioned Visual Value Grounding in MLLMs

2026-04-07 · Zhipin Wang, Christoph Leiter, Christian Frey, Mohamed Hesham Ibrahim Abdalla 외 arxiv

Cultural values are expressed not only through language but also through visual scenes and everyday social practices. Yet existing evaluations of cultural values in language models are almost entirely text-only, leaving …