paper-with-me

Papers

SHARD: Safe and Helpful Alignment via Self-Reframing Distillation

2026-06-14 · Viswonathan Manoranjan, Amogh Gupta, Anvesh Rao Vijjini, Thomas Hofweber, Snigdha Chaturvedi arxiv

Large language models often struggle with sensitive prompts. They may refuse outright, provide generic safety boilerplate, or fail to address the user's legitimate informational needs that can be answered safely. We introduce SHARD, a self-reframing distillation method to improve safe-helpfulness. It first rewrites sensitive prompts to surface benign intent using philosophical guidelines, then reframes its original responses into safe, more helpful ones, and finally fine-tunes the model on its self-reframed responses. Across DNA and the English subset of LINGUASAFE, SHARD improves helpfulness for most model families while preserving safety. It also remains competitive with distillation from a larger teacher model, suggesting that models can internalize safe and helpful behavior elicited from their own. Warning: This paper contains content that may be offensive or harmful.

📄 PDF Abstract BibTeX arXiv:2606.15517

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Simple and Efficient Jailbreak Method Exploiting LLMs' Helpfulness

2025-09-17 · Xuan Luo, Yue Wang, Zefeng He, Geng Tu 외 arxiv

This study reveals a critical safety blind spot in modern LLMs: learning-style queries, which closely resemble ordinary educational questions, can reliably elicit harmful responses. The learning-style queries are constru…

Does "Reasoning" with Large Language Models Improve Recognizing, Generating, and Reframing Unhelpful Thoughts?

2025-03-31 · Yilin Qi, Dong Won Lee, Cynthia Breazeal, Hae Won Park

Cognitive Reframing, a core element of Cognitive Behavioral Therapy (CBT), helps individuals reinterpret negative experiences by finding positive meaning. Recent advances in Large Language Models (LLMs) have demonstrated…

Can Safety Emerge from Weak Supervision? A Systematic Analysis of Small Language Models

2026-03-07 · Punyajoy Saha, Sudipta Halder, Debjyoti Mondal, Subhadarshi Panda arxiv

Safety alignment is critical for deploying large language models (LLMs) in real-world applications, yet most existing approaches rely on large human-annotated datasets and static red-teaming benchmarks that are costly, d…

SPA: Achieving Consensus in LLM Alignment via Self-Priority Optimization

2025-11-09 · Yue Huang, Xiangqi Wang, Xiangliang Zhang arxiv

In high-stakes scenarios-such as self-harm, legal, or medical queries-LLMs must be both trustworthy and helpful. However, these goals often conflict. We propose priority alignment, a new alignment paradigm that enforces …

Training Models to Generate, Recognize, and Reframe Unhelpful Thoughts

2023-07-06 · Mounica Maddela, Megan Ung, Jing Xu, Andrea Madotto 외

Many cognitive approaches to well-being, such as recognizing and reframing unhelpful thoughts, have received considerable empirical support over the past decades, yet still lack truly widespread adoption in self-help for…