paper-with-me

홈 › Papers

AR2: Attention-Guided Repair for the Robustness of CNNs Against Common Corruptions

2025-07-08 · Fuyuan Zhang, Qichen Wang, Jianjun Zhao arxiv

Deep neural networks suffer from significant performance degradation when exposed to common corruptions such as noise, blur, weather, and digital distortions, limiting their reliability in real-world applications. In this paper, we propose AR2 (Attention-Guided Repair for Robustness), a simple yet effective method to enhance the corruption robustness of pretrained CNNs. AR2 operates by explicitly aligning the class activation maps (CAMs) between clean and corrupted images, encouraging the model to maintain consistent attention even under input perturbations. Our approach follows an iterative repair strategy that alternates between CAM-guided refinement and standard fine-tuning, without requiring architectural changes. Extensive experiments show that AR2 consistently outperforms existing state-of-the-art methods in restoring robustness on standard corruption benchmarks (CIFAR-10-C, CIFAR-100-C and ImageNet-C), achieving a favorable balance between accuracy on clean data and corruption robustness. These results demonstrate that AR2 provides a robust and scalable solution for enhancing model reliability in real-world environments with diverse corruptions.

📄 PDF Abstract BibTeX arXiv:2507.06332

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Patch-Fool: Are Vision Transformers Always Robust Against Adversarial Perturbations?

2022-03-16 · ICLR 2022 4 · Yonggan Fu, Shunyao Zhang, Shang Wu, Cheng Wan 외

Vision transformers (ViTs) have recently set off a new wave in neural architecture design thanks to their record-breaking performance in various vision tasks. In parallel, to fulfill the goal of deploying ViTs into real-…

Reveal of Vision Transformers Robustness against Adversarial Attacks

2021-06-07 · Ahmed Aldahdooh, Wassim Hamidouche, Olivier Deforges

The major part of the vanilla vision transformer (ViT) is the attention block that brings the power of mimicking the global context of the input image. For better performance, ViT needs large-scale training data. To over…

Image Classification

Repairing Catastrophic-Neglect in Text-to-Image Diffusion Models via Attention-Guided Feature Enhancement

2024-06-24 · Zhiyuan Chang, Mingyang Li, Junjie Wang, Yi Liu 외

Text-to-Image Diffusion Models (T2I DMs) have garnered significant attention for their ability to generate high-quality images from textual descriptions. However, these models often produce images that do not fully align…

Image Generation

Global Feature Guided Local Pooling

2019-10-01 · ICCV 2019 10 · Takumi Kobayashi

In deep convolutional neural networks (CNNs), local pooling operation is a key building block to effectively downsize feature maps for reducing computation cost as well as increasing robustness against input variation. T…

General Classificationimage-classificationImage Classification

Predict, Reuse, and Repair: Accelerating Dynamic Sparse Attention for Long-Context LLM Decoding

2026-06-29 · Tianyu Wang, Gourav Rattihalli, Aditya Dhakal, Junbo Li 외 arxiv

Dynamic sparse attention (DSA) accelerates long-context LLM decoding by attending to only the top-K KV blocks relevant to each query, but it introduces a serialized selection-to-attention dependency that emerges as a new…