paper-with-me

홈 › Papers

Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding

2025-08-28 · Yuan Xie, Tianshui Chen, Zheng Ge, Lionel Ni arxiv

Long-form video understanding, characterized by long-range temporal dependencies and multiple events, remains a challenge. Existing methods often rely on static reasoning or external visual-language models (VLMs), which face issues like complexity and sub-optimal performance due to the lack of end-to-end training. In this paper, we propose Video-MTR, a reinforced multi-turn reasoning framework designed to enable iterative key video segment selection and question comprehension. Unlike traditional video reasoning pipeline, which generate predictions in a single turn, Video-MTR performs reasoning in multiple turns, selecting video segments progressively based on the evolving understanding of previously processed segments and the current question. This iterative process allows for a more refined and contextually aware analysis of the video. To ensure intermediate reasoning process, we introduce a novel gated bi-level reward system, combining trajectory-level rewards based on answer correctness and turn-level rewards emphasizing frame-query relevance. This system optimizes both video segment selection and question comprehension, eliminating the need for external VLMs and allowing end-to-end training. Extensive experiments on benchmarks like VideoMME, MLVU, and EgoSchema demonstrate that Video-MTR outperforms existing methods in both accuracy and efficiency, advancing the state-of-the-art in long video understanding.

📄 PDF Abstract BibTeX arXiv:2508.20478

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding

2026-02-21 · Houlun Chen, Xin Wang, Guangyao Li, Yuwei Zhou 외 arxiv

Long video understanding is challenging due to rich and complicated multimodal clues in long temporal range.Current methods adopt reasoning to improve the model's ability to analyze complex video clues in long videos via…

Video Grounding

Agentic Reinforced Policy Optimization

2025-07-26 · Guanting Dong, Hangyu Mao, Kai Ma, Licheng Bao 외 arxiv

Large-scale reinforcement learning with verifiable rewards (RLVR) has demonstrated its effectiveness in harnessing the potential of large language models (LLMs) for single-turn reasoning tasks. In realistic reasoning sce…

Reinforcement Learning

SAGE: Training Smart Any-Horizon Agents for Long Video Reasoning with Reinforcement Learning

2025-12-15 · Jitesh Jain, Jialuo Li, Zixian Ma, Jieyu Zhang 외 arxiv

As humans, we are natural any-horizon reasoners, i.e., we can decide whether to iteratively skim long videos or watch short ones in full when necessary for a given task. With this in mind, one would expect video reasonin…

Synthetic Data GenerationReinforcement Learning

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation

2026-07-30 · Xingjian Wu, Junlin Liu, Xingchen Liu, Xuhang Zhu 외 arxiv

Recent advances in post-training Large Language Models (LLMs) increasingly rely on Reinforcement Learning with Verifiable Rewards (RLVR) or On-Policy Self-Distillation (OPSD). While OPSD provides dense, logit-level super…

Reinforcement LearningContrastive Learning

Reinforced Dynamic Reasoning for Conversational Question Generation

2019-07-29 · ACL 2019 7 · Boyuan Pan, Hao Li, Ziyu Yao, Deng Cai 외

This paper investigates a new task named Conversational Question Generation (CQG) which is to generate a question based on a passage and a conversation history (i.e., previous turns of question-answer pairs). CQG is a cr…

DecoderQuestion AnsweringQuestion GenerationQuestion-Generation+1