paper-with-me

홈 › Papers

Multi-modal Feature Fusion with Feature Attention for VATEX Captioning Challenge 2020

2020-06-05 · Ke Lin, Zhuoxin Gan, Li-Wei Wang

This report describes our model for VATEX Captioning Challenge 2020. First, to gather information from multiple domains, we extract motion, appearance, semantic and audio features. Then we design a feature attention module to attend on different feature when decoding. We apply two types of decoders, top-down and X-LAN and ensemble these models to get the final result. The proposed method outperforms official baseline with a significant gap. We achieve 76.0 CIDEr and 50.0 CIDEr on English and Chinese private test set. We rank 2nd on both English and Chinese private test leaderboard.

📄 PDF Abstract BibTeX arXiv:2006.03315

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Tri-attention Fusion Guided Multi-modal Segmentation Network

2021-11-02 · Tongxue Zhou, Su Ruan, Pierre Vera, Stéphane Canu

In the field of multimodal segmentation, the correlation between different modalities can be considered for improving the segmentation results. Considering the correlation between different MR modalities, in this paper, …

Brain Tumor SegmentationSegmentationTumor Segmentation

Attention-guided Multi-step Fusion: A Hierarchical Fusion Network for Multimodal Recommendation

2023-04-24 · Yan Zhou, Jie Guo, Hao Sun, Bin Song 외

The main idea of multimodal recommendation is the rational utilization of the item's multimodal information to improve the recommendation performance. Previous works directly integrate item multimodal features with item …

Contrastive LearningMultimodal Recommendation

Cascaded information enhancement and cross-modal attention feature fusion for multispectral pedestrian detection

2023-02-17 · Yang Yang, Kaixiong Xu, Kaizheng Wang

Multispectral pedestrian detection is a technology designed to detect and locate pedestrians in Color and Thermal images, which has been widely used in automatic driving, video surveillance, etc. So far most available mu…

Pedestrian Detection

MSAF: Multimodal Split Attention Fusion

2020-12-13 · Lang Su, Chuqing Hu, Guofa Li, Dongpu Cao

Multimodal learning mimics the reasoning process of the human multi-sensory system, which is used to perceive the surrounding world. While making a prediction, the human brain tends to relate crucial cues from multiple s…

Action RecognitionEmotion RecognitionMultimodal Emotion RecognitionMultimodal Sentiment Analysis+1

Selective Complementary Feature Fusion and Modal Feature Compression Interaction for Brain Tumor Segmentation

2025-03-20 · Dong Chen, Boyue Zhao, Yi Zhang, Meng Zhao

Efficient modal feature fusion strategy is the key to achieve accurate segmentation of brain glioma. However, due to the specificity of different MRI modes, it is difficult to carry out cross-modal fusion with large diff…

Brain Tumor SegmentationFeature CompressionSpecificityTumor Segmentation