paper-with-me

홈 › Papers

DnA: Denoising Attention for Visual Tasks

2026-06-25 · Ron Campos, Subhajit Maity, Xin Li, Srijan Das, Aritra Dutta arxiv

The softmax activation in multihead attention (MHA) is the de facto standard for attention-based models in visual perception tasks. However, standard softmax can produce noisy attention patterns that dilute relevant features and degrade its performance. In this paper, we propose Denoising Attention or DnA, in which, first, a positive query identifies which image features belong to the correct class, and a negative query identifies closely associated but irrelevant image features. DnA then projects these interactions into two distinct subspaces with larger principal angles, promoting subspace separation and improved discriminability. Using a ViT-B backbone, our proposed DnA achieves an absolute gain of 0.8% on ImageNet-1K compared to the baseline. We further show improvements across multiple visual understanding tasks, including video understanding with video transformers (1.8%) and video LLMs (0.5%). Our extensive empirical analyses justify the design choices involving two interacting subspaces and the denoising effect of DnA.

📄 PDF Abstract BibTeX arXiv:2606.27372

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Two-stage Deep Denoising with Self-guided Noise Attention for Multimodal Medical Images

2025-03-10 · S M A Sharif, Rizwan Ali Naqvi, Woong-Kee Loh

Medical image denoising is considered among the most challenging vision tasks. Despite the real-world implications, existing denoising methods have notable drawbacks as they often generate visual artifacts when applied t…

DenoisingImage DenoisingMedical Image DenoisingSSIM

Learning Medical Image Denoising with Deep Dynamic Residual Attention Network

2020-12-09 · S M A Sharif, Rizwan Ali Naqvi 2, Mithun Biswas

Image denoising performs a prominent role in medical image analysis. In many cases, it can drastically accelerate the diagnostic process by enhancing the perceptual quality of noisy image samples. However, despite the ex…

DenoisingDiagnosticFeature CorrelationImage Denoising+2

Hierarchical Denoising For Multi-Step Visual Reasoning

2026-07-16 · Zezhong Qian, Xiaowei Chi, Chak-Wing Mak, Tianze Zhou 외 arxiv

Video models are evolving into vision foundation models, yet they still lack human-like multi-step reasoning. Streaming autoregressive diffusion models are efficient but limited in reasoning, while bidirectional diffusio…

Visual ReasoningVideo Generation

D$^{3}$ToM: Decider-Guided Dynamic Token Merging for Accelerating Diffusion MLLMs

2025-11-15 · Shuochen Chang, Xiaofeng Zhang, Qingyang Liu, Li Niu arxiv

Diffusion-based multimodal large language models (Diffusion MLLMs) have recently demonstrated impressive non-autoregressive generative capabilities across vision-and-language tasks. However, Diffusion MLLMs exhibit subst…

Reinforced Label Denoising for Weakly-Supervised Audio-Visual Video Parsing

2024-12-27 · Yongbiao Gao, Xiangcheng Sun, Guohua Lv, Deng Yu 외

Audio-visual video parsing (AVVP) aims to recognize audio and visual event labels with precise temporal boundaries, which is quite challenging since audio or visual modality might include only one event label with only t…

Denoising