paper-with-me

Papers

Disentangling Safe and Unsafe Corruptions via Anisotropy and Locality

2025-01-30 · Ramchandran Muthukumar, Ambar Pal, Jeremias Sulam, Rene Vidal

State-of-the-art machine learning systems are vulnerable to small perturbations to their input, where ``small'' is defined according to a threat model that assigns a positive threat to each perturbation. Most prior works define a task-agnostic, isotropic, and global threat, like the $\ell_p$ norm, where the magnitude of the perturbation fully determines the degree of the threat and neither the direction of the attack nor its position in space matter. However, common corruptions in computer vision, such as blur, compression, or occlusions, are not well captured by such threat models. This paper proposes a novel threat model called \texttt{Projected Displacement} (PD) to study robustness beyond existing isotropic and global threat models. The proposed threat model measures the threat of a perturbation via its alignment with \textit{unsafe directions}, defined as directions in the input space along which a perturbation of sufficient magnitude changes the ground truth class label. Unsafe directions are identified locally for each input based on observed training data. In this way, the PD threat model exhibits anisotropy and locality. Experiments on Imagenet-1k data indicate that, for any input, the set of perturbations with small PD threat includes \textit{safe} perturbations of large $\ell_p$ norm that preserve the true label, such as noise, blur and compression, while simultaneously excluding \textit{unsafe} perturbations that alter the true label. Unlike perceptual threat models based on embeddings of large-vision models, the PD threat model can be readily computed for arbitrary classification tasks without pre-training or finetuning. Further additional task annotation such as sensitivity to image regions or concept hierarchies can be easily integrated into the assessment of threat and thus the PD threat model presents practitioners with a flexible, task-driven threat specification.

📄 PDF Abstract BibTeX arXiv:2501.18098

Code (1)

ramcha24/nonisotropic 공식 구현 pytorch

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Disentangling Safe and Unsafe Image Corruptions via Anisotropy and Locality

2025-01-01 · CVPR 2025 1 · Ramchandran Muthukumar, Ambar Pal, Jeremias Sulam, Rene Vidal

State-of-the-art machine learning systems are vulnerable to small perturbations to their input, where _small_ is defined according to a threat model that assigns a positive threat to each perturbation. Most prior wor…

Crash-Consistent Checkpointing for AI Training on macOS/APFS

2025-11-23 · Juha Jeon arxiv

Deep learning training relies on periodic checkpoints to recover from failures, but unsafe checkpoint installation can leave corrupted files on disk. This paper presents an experimental study of checkpoint installation p…

3DEditSafe: Defending 3D Editing Pipelines from Unsafe Generation

2026-05-14 · Nicole Meng, Zheyuan Liu, Meng Jiang, Yingjie Lao arxiv

Recent advances in 3D generative editing, particularly pipelines based on 3D Gaussian Splatting (3DGS), have achieved high-fidelity, multi-view-consistent scene manipulation from text prompts. However, we find that these…

UnsafeChain: Enhancing Reasoning Model Safety via Hard Cases

2025-07-29 · Raj Vardhan Tomar, Preslav Nakov, Yuxia Wang arxiv

As large reasoning models (LRMs) grow more capable, chain-of-thought (CoT) reasoning introduces new safety challenges. Existing SFT-based safety alignment studies dominantly focused on filtering prompts with safe, high-q…

Towards Understanding Unsafe Video Generation

2024-07-17 · Yan Pang, Aiping Xiong, Yang Zhang, Tianhao Wang

Video generation models (VGMs) have demonstrated the capability to synthesize high-quality output. It is important to understand their potential to produce unsafe content, such as violent or terrifying videos. In this wo…

Image GenerationVideo Generation