Attending Self-Attention: A Case Study of Visually Grounded Supervision in Vision-and-Language Transformers
The impressive performances of pre-trained visually grounded language models have motivated a growing body of research investigating what has been learned during the pre-training. As a lot of these models are based on Transformers, several studies on the attention mechanisms used by the models to learn to associate phrases with their visual grounding in the image have been conducted. In this work, we investigate how supervising attention directly to learn visual grounding can affect the behavior of such models. We compare three different methods on attention supervision and their impact on the performances of a state-of-the-art visually grounded language model on two popular vision-and-language tasks.
Code (0)
등록된 구현이 없습니다.
Tasks
Language ModelingLanguage ModellingVisual GroundingSimilar Papers 제목 키워드 기반
Unified Local and Global Attention Interaction Modeling for Vision Transformers
We present a novel method that extends the self-attention mechanism of a vision transformer (ViT) for more accurate object detection across diverse datasets. ViTs show strong capability for image understanding tasks such…
object-detectionObject DetectionAttending to Emotional Narratives
Attention mechanisms in deep neural networks have achieved excellent performance on sequence-prediction tasks. Here, we show that these recently-proposed attention-based mechanisms---in particular, the Transformer with i…
Emotion RecognitionTime SeriesTime Series AnalysisAttend More Times for Image Captioning
Most attention-based image captioning models attend to the image once per word. However, attending once per word is rigid and is easy to miss some information. Attending more times can adjust the attention position, find…
Image CaptioningBridging the Gap: Attending to Discontinuity in Identification of Multiword Expressions
We introduce a new method to tag Multiword Expressions (MWEs) using a linguistically interpretable language-independent deep learning architecture. We specifically target discontinuity, an under-explored aspect that pose…
TAGVisually-Augmented Language Modeling
Human language is grounded on multimodal knowledge including visual knowledge like colors, sizes, and shapes. However, current large-scale pre-trained language models rely on text-only self-supervised training with massi…
Image RetrievalLanguage ModelingLanguage ModellingRetrieval