paper-with-me

홈 › Papers

Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models

2026-05-08 · Yuancheng Wei, Linli Yao, Lei Li, Haojie Zhang, Hao Zhou, Fandong Meng, Xu Sun arxiv

Multimodal reward models have advanced substantially in text and image domains, yet progress in video understanding reward modeling remains severely limited by the lack of robust evaluation benchmarks and high-quality preference data. To address this, we propose a unified framework spanning benchmark design, data construction, and reward model training. We introduce Video Understanding Reward Bench (VURB), a benchmark featuring 2,100 preference pairs with long chain-of-thought reasoning traces (averaging 1,143 tokens) and majority voting evaluation across general, long, and reasoning-oriented video tasks. We further construct Video Understanding Preference Dataset (VUP-35K) via a fully automated pipeline, providing large-scale high-quality supervision for video reward training. Building on the data, we train VideoDRM and VideoGRM, a discriminative and a generative reward model, both achieving state-of-the-art performance on VURB and VideoRewardBench. Further analysis confirms that VUP-35K enhances both reward performance and model reasoning capability, while VideoDRM and VideoGRM yield significant gains under best-of-$N$ test-time scaling.

📄 PDF Abstract BibTeX arXiv:2605.07872

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

2024-10-04 · Wenhao Chai, Enxin Song, Yilun Du, Chenlin Meng 외

Video detailed captioning is a key task which aims to generate comprehensive and coherent textual descriptions of video content, benefiting both video understanding and generation. In this paper, we propose AuroraCap, a …

Image CaptioningVideo Understanding

How Can Objects Help Video-Language Understanding?

2025-04-10 · Zitian Tang, Shijie Wang, Junho Cho, Jaewook Yoo 외

How multimodal large language models (MLLMs) perceive the visual world remains a mystery. To one extreme, object and relation modeling may be implicitly implemented with inductive biases, for example by treating objects …

Image CaptioningObjectQuestion AnsweringVideo Question Answering+1

VIDEOP2R: Video Understanding from Perception to Reasoning

2025-11-14 · Yifan Jiang, Yueying Wang, Rui Zhao, Toufiq Parag 외 arxiv

Reinforcement fine-tuning (RFT), a two-stage framework consisting of supervised fine-tuning (SFT) and reinforcement learning (RL) has shown promising results on improving reasoning ability of large language models (LLMs)…

Reinforcement Learning

A Control-Centric Benchmark for Video Prediction

2023-04-26 · Stephen Tian, Chelsea Finn, Jiajun Wu

Video is a promising source of knowledge for embodied agents to learn models of the world's dynamics. Large deep networks have become increasingly effective at modeling complex video data in a self-supervised manner, as …

PredictionVideo Prediction

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding

2025-08-30 · Zhihong Zhang, Xiaojian Huang, Jin Xu, Zhuodong Luo 외 arxiv

Multimodal reward models (MRMs) play a crucial role in the training, inference, and evaluation of Large Vision Language Models (LVLMs) by assessing response quality. However, existing benchmarks for evaluating MRMs in th…

Reinforcement Learning