paper-with-me

홈 › Papers

Towards Safety and Helpfulness Balanced Responses via Controllable Large Language Models

2024-04-01 · Yi-Lin Tuan, Xilun Chen, Eric Michael Smith, Louis Martin, Soumya Batra, Asli Celikyilmaz, William Yang Wang, Daniel M. Bikel

As large language models (LLMs) become easily accessible nowadays, the trade-off between safety and helpfulness can significantly impact user experience. A model that prioritizes safety will cause users to feel less engaged and assisted while prioritizing helpfulness will potentially cause harm. Possible harms include teaching people how to build a bomb, exposing youth to inappropriate content, and hurting users' mental health. In this work, we propose to balance safety and helpfulness in diverse use cases by controlling both attributes in LLM. We explore training-free and fine-tuning methods that do not require extra human annotations and analyze the challenges of controlling safety and helpfulness in LLMs. Our experiments demonstrate that our method can rewind a learned model and unlock its controllability.

📄 PDF Abstract BibTeX arXiv:2404.01295

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Beyond Static Alignment: Hierarchical Policy Control for LLM Safety via Risk-Aware Chain-of-Thought

2026-02-06 · Jianfeng Si, Lin Sun, Weihong Lin, Xiangzheng Zhang arxiv

Large Language Models (LLMs) face a fundamental safety-helpfulness trade-off due to static, one-size-fits-all safety policies that lack runtime controllabilityxf, making it difficult to tailor responses to diverse applic…

FINEST: Improving LLM Responses to Sensitive Topics Through Fine-Grained Evaluation

2026-03-04 · Juhyun Oh, Nayeon Lee, Chani Jung, Jiho Jin 외 arxiv

Large Language Models (LLMs) often generate overly cautious and vague responses on sensitive topics, sacrificing helpfulness for safety. Existing evaluation frameworks lack systematic methods to identify and address spec…

Equilibrate RLHF: Towards Balancing Helpfulness-Safety Trade-off in Large Language Models

2025-02-17 · Yingshui Tan, Yilei Jiang, Yanshi Li, Jiaheng Liu 외

Fine-tuning large language models (LLMs) based on human preferences, commonly achieved through reinforcement learning from human feedback (RLHF), has been effective in improving their performance. However, maintaining LL…

Safety Alignment

Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment

2024-02-29 · Yiju Guo, Ganqu Cui, Lifan Yuan, Ning Ding 외

Alignment in artificial intelligence pursues the consistency between model responses and human preferences as well as values. In practice, the multifaceted nature of human preferences inadvertently introduces what is kno…

Navigate

SHARD: Safe and Helpful Alignment via Self-Reframing Distillation

2026-06-14 · Viswonathan Manoranjan, Amogh Gupta, Anvesh Rao Vijjini, Thomas Hofweber 외 arxiv

Large language models often struggle with sensitive prompts. They may refuse outright, provide generic safety boilerplate, or fail to address the user's legitimate informational needs that can be answered safely. We intr…