paper-with-me

Papers

Fundamental Safety-Capability Trade-offs in Fine-tuning Large Language Models

2025-03-24 · Pin-Yu Chen, Han Shen, Payel Das, Tianyi Chen

Fine-tuning Large Language Models (LLMs) on some task-specific datasets has been a primary use of LLMs. However, it has been empirically observed that this approach to enhancing capability inevitably compromises safety, a phenomenon also known as the safety-capability trade-off in LLM fine-tuning. This paper presents a theoretical framework for understanding the interplay between safety and capability in two primary safety-aware LLM fine-tuning strategies, providing new insights into the effects of data similarity, context overlap, and alignment loss landscape. Our theoretical results characterize the fundamental limits of the safety-capability trade-off in LLM fine-tuning, which are also validated by numerical experiments.

📄 PDF Abstract BibTeX arXiv:2503.20807

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

What Is the Alignment Tax?

2026-02-09 · Robin Young arxiv

The alignment tax is widely discussed but has not been formally characterized. We provide a geometric theory of the alignment tax in representation space. Under linear representation assumptions, we define the alignment …

Mind the Performance Gap: Capability-Behavior Trade-offs in Feature Steering

2026-02-03 · Eitan Sprejer, Oscar Agustín Stanchi, María Victoria Carro, Denise Alejandra Mester 외 arxiv

Feature steering has emerged as a promising approach for controlling LLM behavior through direct manipulation of internal representations, offering advantages over prompt engineering. However, its practical effectiveness…

Prompt Engineering

Breaking the Safety-Capability Tradeoff: Reinforcement Learning with Verifiable Rewards Maintains Safety Guardrails in LLMs

2025-11-26 · Dongkyu Derek Cho, Huan Song, Arijit Ghosh Chowdhury, Haotian An 외 arxiv

Fine-tuning large language models (LLMs) for downstream tasks typically exhibit a fundamental safety-capability tradeoff, where improving task performance degrades safety alignment even on benign datasets. This degradati…

Reinforcement Learning

SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment

2026-06-01 · Hao Li, Jingkun An, Zijun Song, Pengyu Zhu 외 arxiv

Aligning Large Language Models (LLMs) with human values often degrades their general capabilities, termed the alignment tax. Existing methods mitigate this by balancing dual objectives, which heavily rely on massive gene…

An Adversarial Super-Resolution Remedy for Radar Design Trade-offs

2019-03-04 · Karim Armanious, Sherif Abdulatif, Fady Aziz, Urs Schneider 외

Radar is of vital importance in many fields, such as autonomous driving, safety and surveillance applications. However, it suffers from stringent constraints on its design parametrization leading to multiple trade-offs. …

Autonomous DrivingSuper-Resolution