paper-with-me

홈 › Papers

Extreme Miscalibration and the Illusion of Adversarial Robustness

2024-02-27 · Vyas Raina, Samson Tan, Volkan Cevher, Aditya Rawal, Sheng Zha, George Karypis

Deep learning-based Natural Language Processing (NLP) models are vulnerable to adversarial attacks, where small perturbations can cause a model to misclassify. Adversarial Training (AT) is often used to increase model robustness. However, we have discovered an intriguing phenomenon: deliberately or accidentally miscalibrating models masks gradients in a way that interferes with adversarial attack search methods, giving rise to an apparent increase in robustness. We show that this observed gain in robustness is an illusion of robustness (IOR), and demonstrate how an adversary can perform various forms of test-time temperature calibration to nullify the aforementioned interference and allow the adversarial attack to find adversarial examples. Hence, we urge the NLP community to incorporate test-time temperature scaling into their robustness evaluations to ensure that any observed gains are genuine. Finally, we show how the temperature can be scaled during \textit{training} to improve genuine robustness.

📄 PDF Abstract BibTeX arXiv:2402.17509

Code (0)

등록된 구현이 없습니다.

Tasks

Adversarial AttackAdversarial Robustness

Similar Papers 제목 키워드 기반

Breaking the Illusion: Consensus-Based Generative Mitigation of Adversarial Illusions in Multi-Modal Embeddings

2025-11-26 · Fatemeh Akbarian, Anahita Baninajjar, Yingyi Zhang, Ananth Balashankar 외 arxiv

Multi-modal foundation models align images, text, and other modalities in a shared embedding space but remain vulnerable to adversarial illusions [35], where imperceptible perturbations disrupt cross-modal alignment and …

Synthesizing Visual Illusions Using Generative Adversarial Networks

2019-11-21 · Alexander Gomez-Villa, Adrian Martín, Javier Vazquez-Corral, Jesús Malo 외

Visual illusions are a very useful tool for vision scientists, because they allow them to better probe the limits, thresholds and errors of the visual system. In this work we introduce the first ever framework to generat…

Generative Adversarial Network

Imitation Game for Adversarial Disillusion with Multimodal Generative Chain-of-Thought Role-Play

2025-01-31 · Ching-Chun Chang, Fan-Yun Chen, Shih-Hong Gu, Kai Gao 외

As the cornerstone of artificial intelligence, machine perception confronts a fundamental threat posed by adversarial illusions. These adversarial attacks manifest in two primary forms: deductive illusion, where specific…

Adversarial Robustness in Multi-Task Learning: Promises and Illusions

2021-10-26 · Salah Ghamizi, Maxime Cordy, Mike Papadakis, Yves Le Traon

Vulnerability to adversarial attacks is a well-known weakness of Deep Neural networks. While most of the studies focus on single-task neural networks with computer vision datasets, very little research has considered com…

Adversarial RobustnessMulti-Task Learning

Benchmarking the Robustness of Autonomous Driving to Environmental Illusions: A Lane Perception Perspective

2026-07-07 · Tianyuan Zhang, Xianglong Liu, Aishan Liu, Lu Wang 외 arxiv

Environmental illusions (eg., shadows, reflections, and tire marks) are naturally existing yet overlooked phenomena in real-world driving environments. They can disturb visual perception, leading to misinterpretation of …

Visual Question AnsweringAutonomous DrivingLane Detection