paper-with-me

홈 › Papers

The Rise of Darkness: Safety-Utility Trade-Offs in Role-Playing Dialogue Agents

2025-02-28 · Yihong Tang, Kehai Chen, Xuefeng Bai, ZhengYu Niu, Bo wang, Jie Liu, Min Zhang

Large Language Models (LLMs) have made remarkable advances in role-playing dialogue agents, demonstrating their utility in character simulations. However, it remains challenging for these agents to balance character portrayal utility with content safety because this essential character simulation often comes with the risk of generating unsafe content. To address this issue, we first conduct a systematic exploration of the safety-utility trade-off across multiple LLMs. Our analysis reveals that risk scenarios created by villain characters and user queries (referred to as risk coupling) contribute to this trade-off. Building on this, we propose a novel Adaptive Dynamic Multi-Preference (ADMP) method, which dynamically adjusts safety-utility preferences based on the degree of risk coupling and guides the model to generate responses biased toward utility or safety. We further introduce Coupling Margin Sampling (CMS) into coupling detection to enhance the model's ability to handle high-risk scenarios. Experimental results demonstrate that our approach improves safety metrics while maintaining utility.

📄 PDF Abstract BibTeX arXiv:2502.20757

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Exploring Safety-Utility Trade-Offs in Personalized Language Models

2024-06-17 · Anvesh Rao Vijjini, Somnath Basu Roy Chowdhury, Snigdha Chaturvedi

As large language models (LLMs) become increasingly integrated into daily applications, it is essential to ensure they operate fairly across diverse user demographics. In this work, we show that LLMs suffer from personal…

General Knowledge

Cross-Task Defense: Instruction-Tuning LLMs for Content Safety

2024-05-24 · Yu Fu, Wen Xiao, Jia Chen, Jiachen Li 외

Recent studies reveal that Large Language Models (LLMs) face challenges in balancing safety with utility, particularly when processing long texts for NLP tasks like summarization and translation. Despite defenses against…

Feature-Guided SAE Steering for Refusal-Rate Control using Contrasting Prompts

2025-10-26 · Samaksh Bhargav, Zining Zhu arxiv

Large Language Model (LLM) deployment requires guiding the LLM to recognize and not answer unsafe prompts while complying with safe prompts. Previous methods for achieving this require adjusting model weights along with …

Utility-Fairness Trade-Offs and How to Find Them

2024-04-15 · CVPR 2024 1 · Sepehr Dehdashtian, Bashir Sadeghi, Vishnu Naresh Boddeti

When building classification systems with demographic fairness considerations, there are two objectives to satisfy: 1) maximizing utility for the specific task and 2) ensuring fairness w.r.t. a known demographic attribut…

AttributeFairnessRepresentation Learning

Two Birds, One Projection: Harmonizing Safety and Utility in LVLMs via Inference-time Feature Projection

2026-03-16 · Yewon Han, Yumin Seol, EunGyung Kong, Minsoo Jo 외 arxiv

Existing jailbreak defence frameworks for Large Vision-Language Models often suffer from a safety utility tradeoff, where strengthening safety inadvertently degrades performance on general visual-grounded reasoning tasks…