paper-with-me

홈 › Papers

I Can't Believe It's Not Robust: Catastrophic Collapse of Safety Classifiers under Embedding Drift

2026-03-01 · Subramanyam Sahoo, Vinija Jain, Divya Chaudhary, Aman Chadha arxiv

Instruction tuned reasoning models are increasingly deployed with safety classifiers trained on frozen embeddings, assuming representation stability across model updates. We systematically investigate this assumption and find it fails: normalized perturbations of magnitude $σ=0.02$ (corresponding to $\approx 1^\circ$ angular drift on the embedding sphere) reduce classifier performance from $85\%$ to $50\%$ ROC-AUC. Critically, mean confidence only drops $14\%$, producing dangerous silent failures where $72\%$ of misclassifications occur with high confidence, defeating standard monitoring. We further show that instruction-tuned models exhibit 20$\%$ worse class separability than base models, making aligned systems paradoxically harder to safeguard. Our findings expose a fundamental fragility in production AI safety architectures and challenge the assumption that safety mechanisms transfer across model versions.

📄 PDF Abstract BibTeX arXiv:2603.01297

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models

2026-04-08 · Ismail Hossain, Sai Puppala, Jannatul Ferdaus, Md Jahangir Alam 외 arxiv

A guard model fine-tuned on entirely benign data can lose all safety alignment -- not through adversarial manipulation, but through standard domain specialization. We demonstrate this failure across three purpose-built s…

Position: Model Collapse Does Not Mean What You Think

2025-03-05 · Rylan Schaeffer, Joshua Kazdan, Alvan Caleb Arulandu, Sanmi Koyejo

The proliferation of AI-generated content online has fueled concerns over \emph{model collapse}, a degradation in future generative models' performance when trained on synthetic data generated by earlier models. Industry…

Position

Adversarial amplitude swap towards robust image classifiers

2022-03-14 · Chun Yang Tan, Kazuhiko Kawamoto, Hiroshi Kera

The vulnerability of convolutional neural networks (CNNs) to image perturbations such as common corruptions and adversarial perturbations has recently been investigated from the perspective of frequency. In this study, w…

Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers

2025-10-13 · Wenhan Ma, Hailin Zhang, Liang Zhao, Yifan Song 외 arxiv

Reinforcement learning (RL) has emerged as a crucial approach for enhancing the capabilities of large language models. However, in Mixture-of-Experts (MoE) models, the routing mechanism often introduces instability, even…

Reinforcement Learning

Gen-AI for User Safety: A Survey

2024-11-10 · Akshar Prabhu Desai, Tejasvi Ravi, Mohammad Luqman, Mohit Sharma 외

Machine Learning and data mining techniques (i.e. supervised and unsupervised techniques) are used across domains to detect user safety violations. Examples include classifiers used to detect whether an email is spam or …

Survey