paper-with-me

홈 › Papers

VATEX Captioning Challenge 2019: Multi-modal Information Fusion and Multi-stage Training Strategy for Video Captioning

2019-10-13 · Ziqi Zhang, Yaya Shi, Jiutong Wei, Chunfeng Yuan, Bing Li, Weiming Hu

Multi-modal information is essential to describe what has happened in a video. In this work, we represent videos by various appearance, motion and audio information guided with video topic. By following multi-stage training strategy, our experiments show steady and significant improvement on the VATEX benchmark. This report presents an overview and comparative analysis of our system designed for both Chinese and English tracks on VATEX Captioning Challenge 2019.

📄 PDF Abstract BibTeX arXiv:1910.05752

Code (0)

등록된 구현이 없습니다.

Tasks

Video Captioning

Similar Papers 제목 키워드 기반

Multi-modal Feature Fusion with Feature Attention for VATEX Captioning Challenge 2020

2020-06-05 · Ke Lin, Zhuoxin Gan, Li-Wei Wang

This report describes our model for VATEX Captioning Challenge 2020. First, to gather information from multiple domains, we extract motion, appearance, semantic and audio features. Then we design a feature attention modu…

Vatex Video Captioning Challenge 2020: Multi-View Features and Hybrid Reward Strategies for Video Captioning

2019-10-17 · Xinxin Zhu, Longteng Guo, Peng Yao, Shichen Lu 외

This report describes our solution for the VATEX Captioning Challenge 2020, which requires generating descriptions for the videos in both English and Chinese languages. We identified three crucial factors that improve th…

Video Captioning

Integrating Temporal and Spatial Attentions for VATEX Video Captioning Challenge 2019

2019-10-15 · Shizhe Chen, Yida Zhao, Yuqing Song, Qin Jin 외

This notebook paper presents our model in the VATEX video captioning challenge. In order to capture multi-level aspects in the video, we propose to integrate both temporal and spatial attentions for video captioning. The…

Video Captioning

VATEX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language Research

2019-04-06 · ICCV 2019 10 · Xin Wang, Jiawei Wu, Junkun Chen, Lei LI 외

We present a new large-scale multilingual video description dataset, VATEX, which contains over 41,250 videos and 825,000 captions in both English and Chinese. Among the captions, there are over 206,000 English-Chinese p…

Machine TranslationTranslationVideo CaptioningVideo Description+1

Multi-Modal interpretable automatic video captioning

2024-11-11 · Antoine Hanna-Asaad, Decky Aspandi, Titus Zaharia

Video captioning aims to describe video contents using natural language format that involves understanding and interpreting scenes, actions and events that occurs simultaneously on the view. Current approaches have mainl…

Decision MakingVideo Captioning