paper-with-me

홈 › Papers

VLAB: Enhancing Video Language Pre-training by Feature Adapting and Blending

2023-05-22 · Xingjian He, Sihan Chen, Fan Ma, Zhicheng Huang, Xiaojie Jin, Zikang Liu, Dongmei Fu, Yi Yang, Jing Liu, Jiashi Feng

Large-scale image-text contrastive pre-training models, such as CLIP, have been demonstrated to effectively learn high-quality multimodal representations. However, there is limited research on learning video-text representations for general video multimodal tasks based on these powerful features. Towards this goal, we propose a novel video-text pre-training method dubbed VLAB: Video Language pre-training by feature Adapting and Blending, which transfers CLIP representations to video pre-training tasks and develops unified video multimodal models for a wide range of video-text tasks. Specifically, VLAB is founded on two key strategies: feature adapting and feature blending. In the former, we introduce a new video adapter module to address CLIP's deficiency in modeling temporal information and extend the model's capability to encompass both contrastive and generative tasks. In the latter, we propose an end-to-end training method that further enhances the model's performance by exploiting the complementarity of image and video features. We validate the effectiveness and versatility of VLAB through extensive experiments on highly competitive video multimodal tasks, including video text retrieval, video captioning, and video question answering. Remarkably, VLAB outperforms competing methods significantly and sets new records in video question answering on MSRVTT, MSVD, and TGIF datasets. It achieves an accuracy of 49.6, 61.0, and 79.0, respectively. Codes and models will be released.

📄 PDF Abstract BibTeX arXiv:2305.13167

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringRetrievalText RetrievalTGIF-FrameVideo CaptioningVideo Question AnsweringVideo RetrievalVideo-Text RetrievalVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

Adapter 설명 없음
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation

2025-12-08 · Jiehui Huang, Yuechen Zhang, Xu He, Yuan Gao 외 arxiv

Recent video generation models demonstrate impressive synthesis capabilities but remain limited by single-modality conditioning, constraining their holistic world understanding. This stems from insufficient cross-modal i…

Zero-shot GeneralizationMulti-Task LearningVideo Generation

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

2025-01-21 · Yi Wang, Xinhao Li, Ziang Yan, Yinan He 외

This paper aims to improve the performance of video multimodal large language models (MLLM) via long and rich context (LRC) modeling. As a result, we develop a new version of InternVideo2.5 with a focus on enhancing the …

Object TrackingReferring Expression SegmentationReferring Video Object SegmentationVideo Understanding

LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

2023-11-28 · Yanwei Li, Chengyao Wang, Jiaya Jia

In this work, we present a novel method to tackle the token generation challenge in Vision Language Models (VLMs) for video and image understanding, called LLaMA-VID. Current VLMs, while proficient in tasks like image ca…

Image CaptioningQuestion AnsweringVideo-based Generative Performance BenchmarkingVideo Question Answering+2

Frozen CLIP Models are Efficient Video Learners

2022-08-06 · Ziyi Lin, Shijie Geng, Renrui Zhang, Peng Gao 외

Video recognition has been dominated by the end-to-end learning paradigm -- first initializing a video recognition model with weights of a pretrained image model and then conducting end-to-end training on videos. This en…

Action ClassificationDecoderVideo Recognition

Seeing Dynamic Scene in the Dark: A High-Quality Video Dataset With Mechatronic Alignment

2021-01-01 · ICCV 2021 10 · RuiXing Wang, Xiaogang Xu, Chi-Wing Fu, Jiangbo Lu 외

Low-light video enhancement is an important task. Previous work is mostly trained on paired static images or videos. We compile a new dataset formed by our new strategy that contains high-quality spatially-aligned vi…

Low-Light Image EnhancementVideo Enhancement