paper-with-me

홈 › Papers

Contrastive Video-Language Learning with Fine-grained Frame Sampling

2022-10-10 · Zixu Wang, Yujie Zhong, Yishu Miao, Lin Ma, Lucia Specia

Despite recent progress in video and language representation learning, the weak or sparse correspondence between the two modalities remains a bottleneck in the area. Most video-language models are trained via pair-level loss to predict whether a pair of video and text is aligned. However, even in paired video-text segments, only a subset of the frames are semantically relevant to the corresponding text, with the remainder representing noise; where the ratio of noisy frames is higher for longer videos. We propose FineCo (Fine-grained Contrastive Loss for Frame Sampling), an approach to better learn video and language representations with a fine-grained contrastive objective operating on video frames. It helps distil a video by selecting the frames that are semantically equivalent to the text, improving cross-modal correspondence. Building on the well established VideoCLIP model as a starting point, FineCo achieves state-of-the-art performance on YouCookII, a text-video retrieval benchmark with long videos. FineCo also achieves competitive results on text-video retrieval (MSR-VTT), and video question answering datasets (MSR-VTT QA and MSR-VTT MC) with shorter videos.

📄 PDF Abstract BibTeX arXiv:2210.05039

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringRepresentation LearningRetrievalVideo Question AnsweringVideo Retrieval

Similar Papers 제목 키워드 기반

VideoPerceiver: Enhancing Fine-Grained Temporal Perception in Video Multimodal Large Language Models

2025-11-24 · Fufangchen Zhao, Liao Zhang, Daiqi Shi, Yuanjun Gao 외 arxiv

We propose VideoPerceiver, a novel video multimodal large language model (VMLLM) that enhances fine-grained perception in video understanding, addressing VMLLMs' limited ability to reason about brief actions in short cli…

Reinforcement LearningAction Understanding

ActAlign: Zero-Shot Fine-Grained Video Classification via Language-Guided Sequence Alignment

2025-06-28 · Amir Aghdam, Vincent Tao Hu

We address the task of zero-shot fine-grained video classification, where no video examples or temporal annotations are available for unseen action classes. While contrastive vision-language models such as SigLIP demonst…

Dynamic Time WarpingLarge Language ModelOpen Set Learningtext similarity+2

X-CLIP: End-to-End Multi-grained Contrastive Learning for Video-Text Retrieval

2022-07-15 · Yiwei Ma, Guohai Xu, Xiaoshuai Sun, Ming Yan 외

Video-text retrieval has been a crucial and fundamental task in multi-modal research. The development of video-text retrieval has been considerably promoted by large-scale multi-modal contrastive pre-training, which prim…

Contrastive LearningRetrievalText RetrievalVideo Retrieval+1

Beyond Gloss: A Hand-Centric Framework for Gloss-Free Sign Language Translation

2025-07-31 · Sobhan Asasi, Mohamed Ilyas Lakhal, Ozge Mercanoglu Sincan, Richard Bowden arxiv

Sign Language Translation (SLT) is a challenging task that requires bridging the modality gap between visual and linguistic information while capturing subtle variations in hand shapes and movements. To address these cha…

Sign Language Translation

An Efficient COarse-to-fiNE Alignment Framework @ Ego4D Natural Language Queries Challenge 2022

2022-11-16 · Zhijian Hou, Wanjun Zhong, Lei Ji, Difei Gao 외

This technical report describes the CONE approach for Ego4D Natural Language Queries (NLQ) Challenge in ECCV 2022. We leverage our model CONE, an efficient window-centric COarse-to-fiNE alignment framework. Specifically,…

Contrastive LearningNatural Language Queries