paper-with-me

홈 › Papers

Advancing LLM Safe Alignment with Safety Representation Ranking

2025-05-21 · Tianqi Du, Zeming Wei, Quan Chen, Chenheng Zhang, Yisen Wang

The rapid advancement of large language models (LLMs) has demonstrated milestone success in a variety of tasks, yet their potential for generating harmful content has raised significant safety concerns. Existing safety evaluation approaches typically operate directly on textual responses, overlooking the rich information embedded in the model's internal representations. In this paper, we propose Safety Representation Ranking (SRR), a listwise ranking framework that selects safe responses using hidden states from the LLM itself. SRR encodes both instructions and candidate completions using intermediate transformer representations and ranks candidates via a lightweight similarity-based scorer. Our approach directly leverages internal model states and supervision at the list level to capture subtle safety signals. Experiments across multiple benchmarks show that SRR significantly improves robustness to adversarial prompts. Our code will be available upon publication.

📄 PDF Abstract BibTeX arXiv:2505.15710

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Representation-based Reward Modeling for Efficient Safety Alignment of Large Language Model

2025-03-13 · Qiyuan Deng, Xuefeng Bai, Kehai Chen, YaoWei Wang 외

Reinforcement Learning (RL) algorithms for safety alignment of Large Language Models (LLMs), such as Direct Preference Optimization (DPO), encounter the challenge of distribution shift. Current approaches typically addre…

Language ModelingLanguage ModellingLarge Language ModelReinforcement Learning (RL)+2

Pharmacist: Safety Alignment Data Curation for Large Language Models against Harmful Fine-tuning

2025-10-11 · Guozhi Liu, Qi Mu, Tiansheng Huang, Xinhua Wang 외 arxiv

Harmful fine-tuning issues present significant safety challenges for fine-tuning-as-a-service in large language models. Existing alignment-stage defenses, e.g., Vaccine, Repnoise, Booster, and T-Vaccine, mitigate harmful…

Computational Efficiency

ERPO: Advancing Safety Alignment via Ex-Ante Reasoning Preference Optimization

2025-04-03 · Kehua Feng, Keyan Ding, Jing Yu, MengHan Li 외

Recent advancements in large language models (LLMs) have accelerated progress toward artificial general intelligence, yet their potential to generate harmful content poses critical safety challenges. Existing alignment m…

Safety Alignment

Deliberative Alignment is Deep, but Uncertainty Remains: Inference time safety improvement in reasoning via attribution of unsafe behavior to base model

2026-04-01 · Pankayaraj Pathmanathan, Furong Huang arxiv

While the wide adoption of refusal training in large language models (LLMs) has showcased improvements in model safety, recent works have highlighted shortcomings due to the shallow nature of these alignment methods. To …

SafeNeuron: Neuron-Level Safety Alignment for Large Language Models

2026-02-12 · Zhaoxin Wang, Jiaming Liang, Fengbin Zhu, Weixiang Zhao 외 arxiv

Large language models (LLMs) and multimodal LLMs are typically safety-aligned before release to prevent harmful content generation. However, recent studies show that safety behaviors are concentrated in a small subset of…