paper-with-me

홈 › Papers

Learning Modality Interaction for Temporal Sentence Localization and Event Captioning in Videos

2020-07-28 · ECCV 2020 8 · Shaoxiang Chen, Wenhao Jiang, Wei Liu, Yu-Gang Jiang

Automatically generating sentences to describe events and temporally localizing sentences in a video are two important tasks that bridge language and videos. Recent techniques leverage the multimodal nature of videos by using off-the-shelf features to represent videos, but interactions between modalities are rarely explored. Inspired by the fact that there exist cross-modal interactions in the human brain, we propose a novel method for learning pairwise modality interactions in order to better exploit complementary information for each pair of modalities in videos and thus improve performances on both tasks. We model modality interaction in both the sequence and channel levels in a pairwise fashion, and the pairwise interaction also provides some explainability for the predictions of target tasks. We demonstrate the effectiveness of our method and validate specific design choices through extensive ablation studies. Our method turns out to achieve state-of-the-art performances on four standard benchmark datasets: MSVD and MSR-VTT (event captioning task), and Charades-STA and ActivityNet Captions (temporal sentence localization task).

📄 PDF Abstract BibTeX arXiv:2007.14164

Code (0)

등록된 구현이 없습니다.

Tasks

Sentence

Similar Papers 제목 키워드 기반

Audio-Visual Event Localization in Unconstrained Videos

2018-03-23 · ECCV 2018 9 · Yapeng Tian, Jing Shi, Bochen Li, Zhiyao Duan 외

In this paper, we introduce a novel problem of audio-visual event localization in unconstrained videos. We define an audio-visual event as an event that is both visible and audible in a video segment. We collect an Audio…

audio-visual event localizationTemporal Localization

Dense Audio-Visual Event Localization under Cross-Modal Consistency and Multi-Temporal Granularity Collaboration

2024-12-17 · Ziheng Zhou, Jinxing Zhou, Wei Qian, Shengeng Tang 외

In the field of audio-visual learning, most research tasks focus exclusively on short videos. This paper focuses on the more practical Dense Audio-Visual Event Localization (DAVEL) task, advancing audio-visual scene unde…

audio-visual event localizationaudio-visual learningScene Understanding

Dual Attention Matching for Audio-Visual Event Localization

2019-10-01 · ICCV 2019 10 · Yu Wu, Linchao Zhu, Yan Yan, Yi Yang

In this paper, we investigate the audio-visual event localization problem. This task is to localize a visible and audible event in a video. Previous methods first divide a video into short segments, and then fuse visual …

audio-visual event localization

Weakly Supervised Dense Event Captioning in Videos

2018-12-10 · NeurIPS 2018 12 · Xuguang Duan, Wenbing Huang, Chuang Gan, Jingdong Wang 외

Dense event captioning aims to detect and describe all events of interest contained in a video. Despite the advanced development in this area, existing methods tackle this task by making use of dense temporal annotations…

Sentence

Exploiting Temporal Relationships in Video Moment Localization with Natural Language

2019-08-11 · Songyang Zhang, Jinsong Su, Jiebo Luo

We address the problem of video moment localization with natural language, i.e. localizing a video segment described by a natural language sentence. While most prior work focuses on grounding the query as a whole, tempor…

Sentence