paper-with-me

Papers

Attending Self-Attention: A Case Study of Visually Grounded Supervision in Vision-and-Language Transformers

2021-08-01 · ACL 2021 5 · Jules Samaran, Noa Garcia, Mayu Otani, Chenhui Chu, Yuta Nakashima

The impressive performances of pre-trained visually grounded language models have motivated a growing body of research investigating what has been learned during the pre-training. As a lot of these models are based on Transformers, several studies on the attention mechanisms used by the models to learn to associate phrases with their visual grounding in the image have been conducted. In this work, we investigate how supervising attention directly to learn visual grounding can affect the behavior of such models. We compare three different methods on attention supervision and their impact on the performances of a state-of-the-art visually grounded language model on two popular vision-and-language tasks.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingVisual Grounding

Similar Papers 제목 키워드 기반

Unified Local and Global Attention Interaction Modeling for Vision Transformers

2024-12-25 · Tan Nguyen, Coy D. Heldermon, Corey Toler-Franklin

We present a novel method that extends the self-attention mechanism of a vision transformer (ViT) for more accurate object detection across diverse datasets. ViTs show strong capability for image understanding tasks such…

object-detectionObject Detection

Attending to Emotional Narratives

2019-07-08 · Zhengxuan Wu, Xiyu Zhang, Tan Zhi-Xuan, Jamil Zaki 외

Attention mechanisms in deep neural networks have achieved excellent performance on sequence-prediction tasks. Here, we show that these recently-proposed attention-based mechanisms---in particular, the Transformer with i…

Emotion RecognitionTime SeriesTime Series Analysis

Attend More Times for Image Captioning

2018-12-08 · Jiajun Du, Yu Qin, Hongtao Lu, Yonghua Zhang

Most attention-based image captioning models attend to the image once per word. However, attending once per word is rigid and is easy to miss some information. Attending more times can adjust the attention position, find…

Image Captioning

Bridging the Gap: Attending to Discontinuity in Identification of Multiword Expressions

2019-02-27 · NAACL 2019 6 · Omid Rohanian, Shiva Taslimipoor, Samaneh Kouchaki, Le An Ha 외

We introduce a new method to tag Multiword Expressions (MWEs) using a linguistically interpretable language-independent deep learning architecture. We specifically target discontinuity, an under-explored aspect that pose…

TAG

Visually-Augmented Language Modeling

2022-05-20 · Weizhi Wang, Li Dong, Hao Cheng, Haoyu Song 외

Human language is grounded on multimodal knowledge including visual knowledge like colors, sizes, and shapes. However, current large-scale pre-trained language models rely on text-only self-supervised training with massi…

Image RetrievalLanguage ModelingLanguage ModellingRetrieval