paper-with-me

Papers

Modal-specific Pseudo Query Generation for Video Corpus Moment Retrieval

2022-10-23 · Minjoon Jung, SeongHo Choi, Joochan Kim, Jin-Hwa Kim, Byoung-Tak Zhang

Video corpus moment retrieval (VCMR) is the task to retrieve the most relevant video moment from a large video corpus using a natural language query. For narrative videos, e.g., dramas or movies, the holistic understanding of temporal dynamics and multimodal reasoning is crucial. Previous works have shown promising results; however, they relied on the expensive query annotations for VCMR, i.e., the corresponding moment intervals. To overcome this problem, we propose a self-supervised learning framework: Modal-specific Pseudo Query Generation Network (MPGN). First, MPGN selects candidate temporal moments via subtitle-based moment sampling. Then, it generates pseudo queries exploiting both visual and textual information from the selected temporal moments. Through the multimodal information in the pseudo queries, we show that MPGN successfully learns to localize the video corpus moment without any explicit annotation. We validate the effectiveness of MPGN on the TVR dataset, showing competitive results compared with both supervised models and unsupervised setting models.

📄 PDF Abstract BibTeX arXiv:2210.12617

Code (1)

minjoong507/MPGN 공식 구현 pytorch

Tasks

Moment RetrievalMultimodal ReasoningRetrievalSelf-Supervised LearningVideo Corpus Moment Retrieval

Similar Papers 제목 키워드 기반

Video sentence grounding with temporally global textual knowledge

2024-04-21 · Cai Chen, Runzhong Zhang, Jianjun Gao, Kejun Wu 외

Temporal sentence grounding involves the retrieval of a video moment with a natural language query. Many existing works directly incorporate the given video and temporally localized query for temporal grounding, overlook…

Contrastive LearningRetrievalSentenceTemporal Sentence Grounding

Query-based Video Summarization with Pseudo Label Supervision

2023-07-04 · Jia-Hong Huang, Luka Murn, Marta Mrak, Marcel Worring

Existing datasets for manually labelled query-based video summarization are costly and thus small, limiting the performance of supervised deep video summarization models. Self-supervision can address the data sparsity ch…

Pseudo LabelVideo Summarization

Hybrid-Tower: Fine-grained Pseudo-query Interaction and Generation for Text-to-Video Retrieval

2025-09-05 · Bangxiang Lan, Ruobing Xie, Ruixiang Zhao, Xingwu Sun 외 arxiv

The Text-to-Video Retrieval (T2VR) task aims to retrieve unlabeled videos by textual queries with the same semantic meanings. Recent CLIP-based approaches have explored two frameworks: Two-Tower versus Single-Tower frame…

Video Retrieval

Improving Audio-Visual Video Parsing with Pseudo Visual Labels

2023-03-04 · Jinxing Zhou, Dan Guo, Yiran Zhong, Meng Wang

Audio-Visual Video Parsing is a task to predict the events that occur in video segments for each modality. It often performs in a weakly supervised manner, where only video event labels are provided, i.e., the modalities…

DenoisingPseudo Label

Motion-example-controlled Co-speech Gesture Generation Leveraging Large Language Models

2025-07-27 · Bohong Chen, Yumeng Li, Youyi Zheng, Yao-Xiang Ding 외 arxiv

The automatic generation of controllable co-speech gestures has recently gained growing attention. While existing systems typically achieve gesture control through predefined categorical labels or implicit pseudo-labels …

Gesture Generation