paper-with-me

Papers

Is Safety Standard Same for Everyone? User-Specific Safety Evaluation of Large Language Models

2025-02-20 · Yeonjun In, Wonjoong Kim, Kanghoon Yoon, Sungchul Kim, Mehrab Tanjim, Kibum Kim, Chanyoung Park

As the use of large language model (LLM) agents continues to grow, their safety vulnerabilities have become increasingly evident. Extensive benchmarks evaluate various aspects of LLM safety by defining the safety relying heavily on general standards, overlooking user-specific standards. However, safety standards for LLM may vary based on a user-specific profiles rather than being universally consistent across all users. This raises a critical research question: Do LLM agents act safely when considering user-specific safety standards? Despite its importance for safe LLM use, no benchmark datasets currently exist to evaluate the user-specific safety of LLMs. To address this gap, we introduce U-SAFEBENCH, the first benchmark designed to assess user-specific aspect of LLM safety. Our evaluation of 18 widely used LLMs reveals current LLMs fail to act safely when considering user-specific safety standards, marking a new discovery in this field. To address this vulnerability, we propose a simple remedy based on chain-of-thought, demonstrating its effectiveness in improving user-specific safety. Our benchmark and code are available at https://github.com/yeonjun-in/U-SafeBench.

📄 PDF Abstract BibTeX arXiv:2502.15086

Code (1)

yeonjun-in/u-safebench 공식 구현

Tasks

Language ModelingLanguage ModellingLarge Language Model

Similar Papers 제목 키워드 기반

Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs

2026-05-22 · Qitao Tan, Xiaoying Song, Arman Akbari, Arash Akbari 외 arxiv

Current safety alignment of foundation models largely follows a \emph{one-size-fits-all} paradigm, applying the same refusal policy across users and contexts. As a result, models may refuse requests that are unsafe for g…

Multi-agent systems with CBF-based controllers -- collision avoidance and liveness from instability

2022-07-11 · Mrdjan Jankovic, Mario Santillo, Yan Wang

Assuring system stability is typically a major control design objective. In this paper, we present a system where instability provides a crucial benefit. We consider multi-agent collision avoidance using Control Barrier …

Collision Avoidance

The Chai Platform's AI Safety Framework

2023-06-05 · Xiaoding Lu, Aleksey Korshuk, Zongyi Liu, William Beauchamp

Chai empowers users to create and interact with customized chatbots, offering unique and engaging experiences. Despite the exciting prospects, the work recognizes the inherent challenges of a commitment to modern safety …

Chatbot

LLM Safety for Children

2025-02-18 · Prasanjit Rath, Hari Shrawgi, Parag Agrawal, Sandipan Dandapat

This paper analyzes the safety of Large Language Models (LLMs) in interactions with children below age of 18 years. Despite the transformative applications of LLMs in various aspects of children's lives such as education…

Towards Agile Text Classifiers for Everyone

2023-02-13 · Maximilian Mozes, Jessica Hoffmann, Katrin Tomanek, Muhamed Kouate 외

Text-based safety classifiers are widely used for content moderation and increasingly to tune generative language model behavior - a topic of growing concern for the safety of digital assistants and chatbots. However, di…

Language ModelingLanguage Modellingtext-classificationText Classification