paper-with-me

홈 › Papers

Top-Down Visual Attention from Analysis by Synthesis

2023-03-23 · CVPR 2023 1 · Baifeng Shi, Trevor Darrell, Xin Wang

Current attention algorithms (e.g., self-attention) are stimulus-driven and highlight all the salient objects in an image. However, intelligent agents like humans often guide their attention based on the high-level task at hand, focusing only on task-related objects. This ability of task-guided top-down attention provides task-adaptive representation and helps the model generalize to various tasks. In this paper, we consider top-down attention from a classic Analysis-by-Synthesis (AbS) perspective of vision. Prior work indicates a functional equivalence between visual attention and sparse reconstruction; we show that an AbS visual system that optimizes a similar sparse reconstruction objective modulated by a goal-directed top-down signal naturally simulates top-down attention. We further propose Analysis-by-Synthesis Vision Transformer (AbSViT), which is a top-down modulated ViT model that variationally approximates AbS, and achieves controllable top-down attention. For real-world applications, AbSViT consistently improves over baselines on Vision-Language tasks such as VQA and zero-shot retrieval where language guides the top-down attention. AbSViT can also serve as a general backbone, improving performance on classification, semantic segmentation, and model robustness.

📄 PDF Abstract BibTeX arXiv:2303.13043

Code (1)

bfshi/AbSViT 공식 구현 pytorch

Tasks

RetrievalSemantic SegmentationVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Residual Connection 설명 없음
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…

Similar Papers 제목 키워드 기반

Prototype memory and attention mechanisms for few shot image generation

2021-09-29 · ICLR 2022 4 · Tianqin Li, Zijie Li, Andrew Luo, Harold Rockwell 외

Recent discoveries indicate that the neural codes in the primary visual cortex (V1) of macaque monkeys are complex, diverse and sparse. This leads us to ponder the computational advantages and functional role of these “g…

Image GenerationOnline Clustering

Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion Models

2023-06-29 · NeurIPS 2023 11 · Simian Luo, Chuanhao Yan, Chenxu Hu, Hang Zhao

The Video-to-Audio (V2A) model has recently gained attention for its practical application in generating audio directly from silent videos, particularly in video/film production. However, previous methods in V2A have lim…

Audio Synthesis

NÜWA: Visual Synthesis Pre-training for Neural visUal World creAtion

2021-11-24 · Chenfei Wu, Jian Liang, Lei Ji, Fan Yang 외

This paper presents a unified multimodal pre-trained model called N\"UWA that can generate new or manipulate existing visual data (i.e., images and videos) for various visual synthesis tasks. To cover language, image, an…

DecoderImage GenerationText to Image GenerationText-to-Image Generation+3

Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering

2017-07-25 · CVPR 2018 6 · Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney 외

Top-down visual attention mechanisms have been used extensively in image captioning and visual question answering (VQA) to enable deeper image understanding through fine-grained analysis and even multiple steps of reason…

Image CaptioningVisual Question AnsweringVisual Question Answering (VQA)

GroundingAnomaly: Spatially-Grounded Diffusion for Few-Shot Anomaly Synthesis

2026-04-09 · Yishen Liu, Hongcang Chen, Pengcheng Zhao, Yunfan Bao 외 arxiv

The performance of visual anomaly inspection in industrial quality control is often constrained by the scarcity of real anomalous samples. Consequently, anomaly synthesis techniques have been developed to enlarge trainin…

Anomaly DetectionImage Generation