paper-with-me

홈 › Papers

Dynamic Fusion With Intra- and Inter-Modality Attention Flow for Visual Question Answering

2019-06-01 · CVPR 2019 6 · Peng Gao, Zhengkai Jiang, Haoxuan You, Pan Lu, Steven C. H. Hoi, Xiaogang Wang, Hongsheng Li

Learning effective fusion of multi-modality features is at the heart of visual question answering. We propose a novel method of dynamically fuse multi-modal features with intra- and inter-modality information flow, which alternatively pass dynamic information between and across the visual and language modalities. It can robustly capture the high-level interactions between language and vision domains, thus significantly improves the performance of visual question answering. We also show that, the proposed dynamic intra modality attention flow conditioned on the other modality can dynamically modulate the intra-modality attention of the current modality, which is vital for multimodality feature fusion. Experimental evaluations on the VQA 2.0 dataset show that the proposed method achieves the state-of-the-art VQA performance. Extensive ablation studies are carried out for the comprehensive analysis of the proposed method.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Dynamic Fusion with Intra- and Inter- Modality Attention Flow for Visual Question Answering

2018-12-13 · Gao Peng, Zhengkai Jiang, Haoxuan You, Pan Lu 외

Learning effective fusion of multi-modality features is at the heart of visual question answering. We propose a novel method of dynamically fusing multi-modal features with intra- and inter-modality information flow, whi…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Multimodal Hyperspectral Image Classification via Interconnected Fusion

2023-04-02 · Lu Huo, Jiahao Xia, Leijie Zhang, Haimin Zhang 외

Existing multiple modality fusion methods, such as concatenation, summation, and encoder-decoder-based fusion, have recently been employed to combine modality characteristics of Hyperspectral Image (HSI) and Light Detect…

ClassificationDecoderHyperspectral Image Classificationimage-classification+1

DM$^2$S$^2$: Deep Multi-Modal Sequence Sets with Hierarchical Modality Attention

2022-09-07 · Shunsuke Kitada, Yuki Iwazaki, Riku Togashi, Hitoshi Iyatomi

There is increasing interest in the use of multimodal data in various web applications, such as digital advertising and e-commerce. Typical methods for extracting important information from multimodal data rely on a mid-…

DF^2AM: Dual-level Feature Fusion and Affinity Modeling for RGB-Infrared Cross-modality Person Re-identification

2021-04-01 · Junhui Yin, Zhanyu Ma, Jiyang Xie, Shibo Nie 외

RGB-infrared person re-identification is a challenging task due to the intra-class variations and cross-modality discrepancy. Existing works mainly focus on learning modality-shared global representations by aligning ima…

Cross-Modality Person Re-identificationPerson Re-Identification

IIANet: An Intra- and Inter-Modality Attention Network for Audio-Visual Speech Separation

2023-08-16 · Kai Li, Runxuan Yang, Fuchun Sun, Xiaolin Hu

Recent research has made significant progress in designing fusion modules for audio-visual speech separation. However, they predominantly focus on multi-modal fusion at a single temporal scale of auditory and visual feat…

Speech Separation