Modal-specific Pseudo Query Generation for Video Corpus Moment Retrieval
Video corpus moment retrieval (VCMR) is the task to retrieve the most relevant video moment from a large video corpus using a natural language query. For narrative videos, e.g., dramas or movies, the holistic understanding of temporal dynamics and multimodal reasoning is crucial. Previous works have shown promising results; however, they relied on the expensive query annotations for VCMR, i.e., the corresponding moment intervals. To overcome this problem, we propose a self-supervised learning framework: Modal-specific Pseudo Query Generation Network (MPGN). First, MPGN selects candidate temporal moments via subtitle-based moment sampling. Then, it generates pseudo queries exploiting both visual and textual information from the selected temporal moments. Through the multimodal information in the pseudo queries, we show that MPGN successfully learns to localize the video corpus moment without any explicit annotation. We validate the effectiveness of MPGN on the TVR dataset, showing competitive results compared with both supervised models and unsupervised setting models.
Code (1)
Tasks
Moment RetrievalMultimodal ReasoningRetrievalSelf-Supervised LearningVideo Corpus Moment RetrievalSimilar Papers 제목 키워드 기반
Video sentence grounding with temporally global textual knowledge
Temporal sentence grounding involves the retrieval of a video moment with a natural language query. Many existing works directly incorporate the given video and temporally localized query for temporal grounding, overlook…
Contrastive LearningRetrievalSentenceTemporal Sentence GroundingQuery-based Video Summarization with Pseudo Label Supervision
Existing datasets for manually labelled query-based video summarization are costly and thus small, limiting the performance of supervised deep video summarization models. Self-supervision can address the data sparsity ch…
Pseudo LabelVideo SummarizationHybrid-Tower: Fine-grained Pseudo-query Interaction and Generation for Text-to-Video Retrieval
The Text-to-Video Retrieval (T2VR) task aims to retrieve unlabeled videos by textual queries with the same semantic meanings. Recent CLIP-based approaches have explored two frameworks: Two-Tower versus Single-Tower frame…
Video RetrievalImproving Audio-Visual Video Parsing with Pseudo Visual Labels
Audio-Visual Video Parsing is a task to predict the events that occur in video segments for each modality. It often performs in a weakly supervised manner, where only video event labels are provided, i.e., the modalities…
DenoisingPseudo LabelMotion-example-controlled Co-speech Gesture Generation Leveraging Large Language Models
The automatic generation of controllable co-speech gestures has recently gained growing attention. While existing systems typically achieve gesture control through predefined categorical labels or implicit pseudo-labels …
Gesture Generation