paper-with-me

홈 › Papers

Safety Measurements for Fine-tuned LLMs Should be Grounded in Capability

2026-06-02 · Krishnapriya Vishnubhotla, Hillary Dawkins, Isar Nejadgholi, Svetlana Kiritchenko arxiv

Adapting foundation large language models to a user's task or preferred style through fine-tuning can result in compromising the model's safety. Previous works examined the effects of fine-tuning on model safety in limited and seemingly random experimental settings. We argue that anchoring fine-tuning to a specific capability goal is essential for avoiding arbitrary empirical choices, allowing us to draw meaningful conclusions about safety impacts, and to compare mitigation methods on a consistent basis. We conduct a multi-dimensional evaluation of the effects of fine-tuning on model behavior by focusing on capability as well as safety. Our results surface important issues that (1) fine-tuned models can produce incoherent generations in response to safety prompts, (2) automated safety judgments are unreliable for such incoherent outputs, and (3) the conclusions about the effects of fine-tuning can change depending on the choice of safety benchmark as well as the safety evaluator.

📄 PDF Abstract BibTeX arXiv:2606.03648

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Locking Down the Finetuned LLMs Safety

2024-10-14 · Minjun Zhu, Linyi Yang, Yifan Wei, Ningyu Zhang 외

Fine-tuning large language models (LLMs) on additional datasets is often necessary to optimize them for specific downstream tasks. However, existing safety alignment measures, which restrict harmful behavior during infer…

Safety Alignment

Emergent alignment and the projectability of ethical personas

2026-06-08 · Guillermo Del Pinal, Youngchan Lee, Calum McNamara, Alejandro Perez Carballo arxiv

Work on `emergent misalignment' shows that finetuning LLMs on narrow tasks can induce broadly misaligned behavior. This supports the `persona selection' (PSM) hypothesis: during pre-training, LLMs learn to simulate diffe…

SafeCOMM: What about Safety Alignment in Fine-Tuned Telecom Large Language Models?

2025-05-29 · Aladin Djuhera, Swanand Ravindra Kadhe, Farhan Ahmed, Syed Zawad 외

Fine-tuning large language models (LLMs) for telecom tasks and datasets is a common practice to adapt general-purpose models to the telecom domain. However, little attention has been paid to how this process may compromi…

DiagnosticRed TeamingSafety Alignment

Separate the Wheat from the Chaff: A Post-Hoc Approach to Safety Re-Alignment for Fine-Tuned Language Models

2024-12-15 · Di wu, Xin Lu, Yanyan Zhao, Bing Qin

Although large language models (LLMs) achieve effective safety alignment at the time of release, they still face various safety challenges. A key issue is that fine-tuning often compromises the safety alignment of LLMs. …

Safety Alignment

Safeguard Fine-Tuned LLMs Through Pre- and Post-Tuning Model Merging

2024-12-27 · Hua Farn, Hsuan Su, Shachi H Kumar, Saurav Sahay 외

Fine-tuning large language models (LLMs) for downstream tasks is a widely adopted approach, but it often leads to safety degradation in safety-aligned LLMs. Currently, many solutions address this issue by incorporating a…