paper-with-me

홈 › Papers

Mask to reconstruct: Cooperative Semantics Completion for Video-text Retrieval

2023-05-13 · Han Fang, Zhifei Yang, Xianghao Zang, Chao Ban, Hao Sun

Recently, masked video modeling has been widely explored and significantly improved the model's understanding ability of visual regions at a local level. However, existing methods usually adopt random masking and follow the same reconstruction paradigm to complete the masked regions, which do not leverage the correlations between cross-modal content. In this paper, we present Mask for Semantics Completion (MASCOT) based on semantic-based masked modeling. Specifically, after applying attention-based video masking to generate high-informed and low-informed masks, we propose Informed Semantics Completion to recover masked semantics information. The recovery mechanism is achieved by aligning the masked content with the unmasked visual regions and corresponding textual context, which makes the model capture more text-related details at a patch level. Additionally, we shift the emphasis of reconstruction from irrelevant backgrounds to discriminative parts to ignore regions with low-informed masks. Furthermore, we design dual-mask co-learning to incorporate video cues under different masks and learn more aligned video representation. Our MASCOT performs state-of-the-art performance on four major text-video retrieval benchmarks, including MSR-VTT, LSMDC, ActivityNet, and DiDeMo. Extensive ablation studies demonstrate the effectiveness of the proposed schemes.

📄 PDF Abstract BibTeX arXiv:2305.07910

Code (0)

등록된 구현이 없습니다.

Tasks

RetrievalText RetrievalVideo RetrievalVideo-Text Retrieval

Similar Papers 제목 키워드 기반

Global and Local Semantic Completion Learning for Vision-Language Pre-training

2023-06-12 · Rong-Cheng Tu, Yatai Ji, Jie Jiang, Weijie Kong 외

Cross-modal alignment plays a crucial role in vision-language pre-training (VLP) models, enabling them to capture meaningful associations across different modalities. For this purpose, numerous masked modeling tasks have…

cross-modal alignmentImage-text RetrievalLanguage ModellingMasked Language Modeling+5

Image Completion via Dual-path Cooperative Filtering

2023-04-30 · Pourya Shamsolmoali, Masoumeh Zareapoor, Eric Granger

Given the recent advances with image-generating algorithms, deep image completion methods have made significant progress. However, state-of-art methods typically provide poor cross-scene generalization, and generated mas…

Seeing What You Miss: Vision-Language Pre-training with Semantic Completion Learning

2022-11-24 · CVPR 2023 1 · Yatai Ji, RongCheng Tu, Jie Jiang, Weijie Kong 외

Cross-modal alignment is essential for vision-language pre-training (VLP) models to learn the correct corresponding information across different modalities. For this purpose, inspired by the success of masked language mo…

cross-modal alignmentImage-text RetrievalLanguage ModelingLanguage Modelling+8

PointSSC: A Cooperative Vehicle-Infrastructure Point Cloud Benchmark for Semantic Scene Completion

2023-09-22 · Yuxiang Yan, Boda Liu, Jianfei Ai, Qinbu Li 외

Semantic Scene Completion (SSC) aims to jointly generate space occupancies and semantic labels for complex 3D scenes. Most existing SSC models focus on volumetric representations, which are memory-inefficient for large o…

Point Cloud CompletionSegmentation

MILES: Visual BERT Pre-training with Injected Language Semantics for Video-text Retrieval

2022-04-26 · Yuying Ge, Yixiao Ge, Xihui Liu, Alex Jinpeng Wang 외

Dominant pre-training work for video-text retrieval mainly adopt the "dual-encoder" architectures to enable efficient retrieval, where two separate encoders are used to contrast global video and text representations, but…

Action RecognitionRetrievalText RetrievalText to Video Retrieval+5