paper-with-me

홈 › Papers

Dynamic Fusion with Intra- and Inter- Modality Attention Flow for Visual Question Answering

2018-12-13 · Gao Peng, Zhengkai Jiang, Haoxuan You, Pan Lu, Steven Hoi, Xiaogang Wang, Hongsheng Li

Learning effective fusion of multi-modality features is at the heart of visual question answering. We propose a novel method of dynamically fusing multi-modal features with intra- and inter-modality information flow, which alternatively pass dynamic information between and across the visual and language modalities. It can robustly capture the high-level interactions between language and vision domains, thus significantly improves the performance of visual question answering. We also show that the proposed dynamic intra-modality attention flow conditioned on the other modality can dynamically modulate the intra-modality attention of the target modality, which is vital for multimodality feature fusion. Experimental evaluations on the VQA 2.0 dataset show that the proposed method achieves state-of-the-art VQA performance. Extensive ablation studies are carried out for the comprehensive analysis of the proposed method.

📄 PDF Abstract BibTeX arXiv:1812.05252

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Dynamic Fusion With Intra- and Inter-Modality Attention Flow for Visual Question Answering

2019-06-01 · CVPR 2019 6 · Peng Gao, Zhengkai Jiang, Haoxuan You, Pan Lu 외

Learning effective fusion of multi-modality features is at the heart of visual question answering. We propose a novel method of dynamically fuse multi-modal features with intra- and inter-modality information flow, which…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Multimodal Hyperspectral Image Classification via Interconnected Fusion

2023-04-02 · Lu Huo, Jiahao Xia, Leijie Zhang, Haimin Zhang 외

Existing multiple modality fusion methods, such as concatenation, summation, and encoder-decoder-based fusion, have recently been employed to combine modality characteristics of Hyperspectral Image (HSI) and Light Detect…

ClassificationDecoderHyperspectral Image Classificationimage-classification+1

DM$^2$S$^2$: Deep Multi-Modal Sequence Sets with Hierarchical Modality Attention

2022-09-07 · Shunsuke Kitada, Yuki Iwazaki, Riku Togashi, Hitoshi Iyatomi

There is increasing interest in the use of multimodal data in various web applications, such as digital advertising and e-commerce. Typical methods for extracting important information from multimodal data rely on a mid-…

DF^2AM: Dual-level Feature Fusion and Affinity Modeling for RGB-Infrared Cross-modality Person Re-identification

2021-04-01 · Junhui Yin, Zhanyu Ma, Jiyang Xie, Shibo Nie 외

RGB-infrared person re-identification is a challenging task due to the intra-class variations and cross-modality discrepancy. Existing works mainly focus on learning modality-shared global representations by aligning ima…

Cross-Modality Person Re-identificationPerson Re-Identification

IIANet: An Intra- and Inter-Modality Attention Network for Audio-Visual Speech Separation

2023-08-16 · Kai Li, Runxuan Yang, Fuchun Sun, Xiaolin Hu

Recent research has made significant progress in designing fusion modules for audio-visual speech separation. However, they predominantly focus on multi-modal fusion at a single temporal scale of auditory and visual feat…

Speech Separation