paper-with-me

Papers

Modulated Self-attention Convolutional Network for VQA

2019-10-08 · Jean-Benoit Delbrouck, Antoine Maiorca, Nathan Hubens, Stéphane Dupont

As new data-sets for real-world visual reasoning and compositional question answering are emerging, it might be needed to use the visual feature extraction as a end-to-end process during training. This small contribution aims to suggest new ideas to improve the visual processing of traditional convolutional network for visual question answering (VQA). In this paper, we propose to modulate by a linguistic input a CNN augmented with self-attention. We show encouraging relative improvements for future research in this direction.

📄 PDF Abstract BibTeX arXiv:1910.03343

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)Visual Reasoning

Similar Papers 제목 키워드 기반

Cross-modulated Attention Transformer for RGBT Tracking

2024-08-05 · Yun Xiao, jiacong Zhao, Andong Lu, Chenglong Li 외

Existing Transformer-based RGBT trackers achieve remarkable performance benefits by leveraging self-attention to extract uni-modal features and cross-attention to enhance multi-modal feature interaction and template-sear…

Rgb-T Tracking

Spectral-Adaptive Modulation Networks for Visual Perception

2025-03-31 · Guhnoo Yun, Juhan Yoo, Kijung Kim, Jeongho Lee 외

Recent studies have shown that 2D convolution and self-attention exhibit distinct spectral behaviors, and optimizing their spectral properties can enhance vision model performance. However, theoretical analyses remain li…

object-detectionObject DetectionSemantic Segmentation

ssToken: Self-modulated and Semantic-aware Token Selection for LLM Fine-tuning

2025-10-21 · Xiaohan Qin, Xiaoxing Wang, Ning Liao, Cancheng Zhang 외 arxiv

Data quality plays a critical role in enhancing supervised fine-tuning (SFT) for large language models (LLMs), and token-level data selection has emerged as a promising direction for its fine-grained nature. Despite thei…

VTAMIQ: Transformers for Attention Modulated Image Quality Assessment

2021-10-04 · Andrei Chubarau, James Clark

Following the major successes of self-attention and Transformers for image analysis, we investigate the use of such attention mechanisms in the context of Image Quality Assessment (IQA) and propose a novel full-reference…

Image Quality Assessment

FAME: Fairness-aware Attention-modulated Video Editing

2025-10-27 · Zhangkai Wu, Xuhui Fan, Zhongyuan Xie, Kaize Shi 외 arxiv

Training-free video editing (VE) models tend to fall back on gender stereotypes when rendering profession-related prompts. We propose \textbf{FAME} for \textit{Fairness-aware Attention-modulated Video Editing} that mitig…