paper-with-me

Papers

Unhackable Temporal Rewarding for Scalable Video MLLMs

2025-02-17 · En Yu, Kangheng Lin, Liang Zhao, Yana Wei, Zining Zhu, Haoran Wei, Jianjian Sun, Zheng Ge, Xiangyu Zhang, Jingyu Wang, Wenbing Tao

In the pursuit of superior video-processing MLLMs, we have encountered a perplexing paradox: the "anti-scaling law", where more data and larger models lead to worse performance. This study unmasks the culprit: "temporal hacking", a phenomenon where models shortcut by fixating on select frames, missing the full video narrative. In this work, we systematically establish a comprehensive theory of temporal hacking, defining it from a reinforcement learning perspective, introducing the Temporal Perplexity (TPL) score to assess this misalignment, and proposing the Unhackable Temporal Rewarding (UTR) framework to mitigate the temporal hacking. Both theoretically and empirically, TPL proves to be a reliable indicator of temporal modeling quality, correlating strongly with frame activation patterns. Extensive experiments reveal that UTR not only counters temporal hacking but significantly elevates video comprehension capabilities. This work not only advances video-AI systems but also illuminates the critical importance of aligning proxy rewards with true objectives in MLLM development.

📄 PDF Abstract BibTeX arXiv:2502.12081

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Incentivizing Temporal-Awareness in Egocentric Video Understanding Models

2026-03-28 · Zhiyang Xu, Tian Qin, Bowen Jin, Zhengfeng Lai 외 arxiv

Multimodal large language models (MLLMs) have recently shown strong performance in visual understanding, yet they often lack temporal awareness, particularly in egocentric settings where reasoning depends on the correct …

Reinforcement Learning

Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoning

2026-06-26 · Hohin Kwan, Hongyu Li, Ray Zhang, Manyuan Zhang 외 arxiv

Recent interest in multimodal large language models (MLLMs) raises a central question: can they reason over dynamic visual evidence rather than merely recognize objects or events in individual frames? This ability, which…

Logical Reasoning

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding

2026-01-03 · Hyeonjeong Ha, Jinjin Ge, Bo Feng, Kaixin Ma 외 arxiv

Multimodal large language models (MLLMs) have achieved impressive progress in vision-language reasoning, yet their ability to understand temporally unfolding narratives in videos remains underexplored. True narrative und…

Defining and Characterizing Reward Hacking

2022-09-27 · Joar Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, David Krueger

We provide the first formal definition of reward hacking, a phenomenon where optimizing an imperfect proxy reward function, $\mathcal{\tilde{R}}$, leads to poor performance according to the true reward function, $\mathca…

Spatial Preference Rewarding for MLLMs Spatial Understanding

2025-10-16 · Han Qiu, Peng Gao, Lewei Lu, Xiaoqin Zhang 외 arxiv

Multimodal large language models~(MLLMs) have demonstrated promising spatial understanding capabilities, such as referencing and grounding object descriptions. Despite their successes, MLLMs still fall short in fine-grai…

Object Localization