paper-with-me

홈 › Papers

Self-Mined Hardness for Safety Fine-Tuning

2026-05-04 · Prakhar Gupta, Garv Shah, Donghua Zhang arxiv

Safety fine-tuning of language models typically requires a curated adversarial dataset. We take a different approach: score each candidate prompt's difficulty by how often the target model's own rollouts are judged harmful, then fine-tune on the hardest prompts paired with the model's own non-jailbroken rollouts. On Llama-3-8B-Instruct and Llama-3.2-3B-Instruct, this approach cuts the WildJailbreak attack success rate from 11.5% and 20.1% down to 1-3%, but pushes refusal on jailbreak-shaped benign prompts from 14-22% to 74-94%. Interleaving the same hard prompts 1:1 with adversarially-framed benign prompts (prompts that look like jailbreaks but have benign intent) cuts that refusal back down to 30-51% on 8B and 52-72% on 3B, at a cost of 2-6 percentage points of attack success rate. Within the mixed regime, training on the hardest half of the eligible pool rather than a random half cuts the remaining ASR by 35-50% (about 3 percentage points) on both models.

📄 PDF Abstract BibTeX arXiv:2605.03226

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Safety Measurements for Fine-tuned LLMs Should be Grounded in Capability

2026-06-02 · Krishnapriya Vishnubhotla, Hillary Dawkins, Isar Nejadgholi, Svetlana Kiritchenko arxiv

Adapting foundation large language models to a user's task or preferred style through fine-tuning can result in compromising the model's safety. Previous works examined the effects of fine-tuning on model safety in limit…

Improving Self-supervised Learning with Hardness-aware Dynamic Curriculum Learning: An Application to Digital Pathology

2021-08-16 · Chetan L Srinidhi, Anne L Martel

Self-supervised learning (SSL) has recently shown tremendous potential to learn generic visual representations useful for many image analysis tasks. Despite their notable success, the existing SSL methods fail to general…

Self-Supervised Learning

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning

2025-06-18 · Gabrel J. Perin, Runjin Chen, Xuxi Chen, Nina S. T. Hirata 외

Large Language Models (LLMs) have become indispensable in real-world applications. However, their widespread adoption raises significant safety concerns, particularly in responding to socially harmful questions. Despite …

Attribute

HAIR: Hardness-Aware Inverse Reinforcement Learning with Introspective Reasoning for LLM Alignment

2025-03-23 · Ruoxi Cheng, Haoxuan Ma, Weixin Wang

The alignment of large language models (LLMs) with human values remains critical yet hindered by four key challenges: (1) scarcity of balanced safety datasets, (2) alignment tax, (3) vulnerability to jailbreak attacks du…

Learning and Forgetting Unsafe Examples in Large Language Models

2023-12-20 · Jiachen Zhao, Zhun Deng, David Madras, James Zou 외

As the number of large language models (LLMs) released to the public grows, there is a pressing need to understand the safety implications associated with these models learning from third-party custom finetuning data. We…