paper-with-me

홈 › Papers

MVP: Enhancing Video Large Language Models via Self-supervised Masked Video Prediction

2026-01-07 · Xiaokun Sun, Zezhong Wu, Zewen Ding, Linli Xu arxiv

Reinforcement learning based post-training paradigms for Video Large Language Models (VideoLLMs) have achieved significant success by optimizing for visual-semantic tasks such as captioning or VideoQA. However, while these approaches effectively enhance perception abilities, they primarily target holistic content understanding, often lacking explicit supervision for intrinsic temporal coherence and inter-frame correlations. This tendency limits the models' ability to capture intricate dynamics and fine-grained visual causality. To explicitly bridge this gap, we propose a novel post-training objective: Masked Video Prediction (MVP). By requiring the model to reconstruct a masked continuous segment from a set of challenging distractors, MVP forces the model to attend to the sequential logic and temporal context of events. To support scalable training, we introduce a scalable data synthesis pipeline capable of transforming arbitrary video corpora into MVP training samples, and further employ Group Relative Policy Optimization (GRPO) with a fine-grained reward function to enhance the model's understanding of video context and temporal properties. Comprehensive evaluations demonstrate that MVP enhances video reasoning capabilities by directly reinforcing temporal reasoning and causal understanding.

📄 PDF Abstract BibTeX arXiv:2601.03781

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningVideo Prediction

Similar Papers 제목 키워드 기반

Learning to Track Instance from Single Nature Language Description

2026-05-08 · Yaozong Zheng, Bineng Zhong, Qihua Liang, Shuimu Zeng 외 arxiv

How to achieve vision-language (VL) tracking using natural language descriptions from a video sequence \textbf{without relying on any bounding-box ground truth}? In this work, we achieve this goal by tackling \textit{sel…

Self-Supervised Learning

NimbleD: Enhancing Self-supervised Monocular Depth Estimation with Pseudo-labels and Large-scale Video Pre-training

2024-08-26 · Albert Luginov, Muhammad Shahzad

We introduce NimbleD, an efficient self-supervised monocular depth estimation learning framework that incorporates supervision from pseudo-labels generated by a large vision model. This framework does not require camera …

Depth EstimationMonocular Depth Estimation

SRVAU-R1: Enhancing Video Anomaly Understanding via Reflection-Aware Learning

2026-02-01 · Zihao Zhao, Shengting Cao, Muchao Ye arxiv

Multi-modal large language models (MLLMs) have demonstrated significant progress in reasoning capabilities and shown promising effectiveness in video anomaly understanding (VAU) tasks. However, existing MLLM-based approa…

Self-Supervised Learning of Deviation in Latent Representation for Co-speech Gesture Video Generation

2024-09-26 · Huan Yang, Jiahui Chen, Chaofan Ding, Runhua Shi 외

Gestures are pivotal in enhancing co-speech communication. While recent works have mostly focused on point-level motion transformation or fully supervised motion representations through data-driven approaches, we explore…

Self-Supervised LearningSSIMVideo Generation

Self-supervised Learning for Semi-supervised Temporal Language Grounding

2021-09-23 · Fan Luo, Shaoxiang Chen, Jingjing Chen, Zuxuan Wu 외

Given a text description, Temporal Language Grounding (TLG) aims to localize temporal boundaries of the segments that contain the specified semantics in an untrimmed video. TLG is inherently a challenging task, as it req…

Contrastive LearningPseudo LabelSelf-Supervised LearningSentence