paper-with-me

홈 › Papers

VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos

2026-02-08 · Wenqi Liu, Yunxiao Wang, Shijie Ma, Meng Liu, Qile Su, Tianke Zhang, Haonan Fan, Changyi Liu, Kaiyu Jiang, Jiankang Chen, Kaiyu Tang, Bin Wen, Fan Yang, Tingting Gao, Han Li, Yinwei Wei, Xuemeng Song arxiv

In long-video understanding, conventional uniform frame sampling often fails to capture key visual evidence, leading to degraded performance and increased hallucinations. To address this, recent agentic thinking-with-videos paradigms have emerged, adopting a localize-clip-answer pipeline in which the model actively identifies relevant video segments, performs dense sampling within those clips, and then produces answers. However, existing methods remain inefficient, suffer from weak localization, and adhere to rigid workflows. To solve these issues, we propose VideoTemp-o3, a unified agentic thinking-with-videos framework that jointly models video grounding and question answering. VideoTemp-o3 exhibits strong localization capability, supports on-demand clipping, and can refine inaccurate localizations. Specifically, in the supervised fine-tuning stage, we design a unified masking mechanism that encourages exploration while preventing noise. For reinforcement learning, we introduce dedicated rewards to mitigate reward hacking. Besides, from the data perspective, we develop an effective pipeline to construct high-quality long video grounded QA data, along with a corresponding benchmark for systematic evaluation across various video durations. Experimental results demonstrate that our method achieves remarkable performance on both long video understanding and grounding.

📄 PDF Abstract BibTeX arXiv:2602.07801

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningQuestion AnsweringVideo Grounding

Similar Papers 제목 키워드 기반

Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

2024-10-04 · Haibo Wang, Zhiyang Xu, Yu Cheng, Shizhe Diao 외

Video Large Language Models (Video-LLMs) have demonstrated remarkable capabilities in coarse-grained video understanding, however, they struggle with fine-grained temporal grounding. In this paper, we introduce Grounded-…

Dense Video CaptioningSentenceTemporal Sentence GroundingVideo Captioning+1

Temporal Preference Optimization for Long-Form Video Understanding

2025-01-23 · Rui Li, Xiaohan Wang, Yuhui Zhang, Zeyu Wang 외

Despite significant advancements in video large multimodal models (video-LMMs), achieving effective temporal grounding in long-form videos remains a challenge for existing models. To address this limitation, we propose T…

FormMMEVideo MMEVideo Understanding

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding

2025-05-27 · Fuwen Luo, Shengfeng Lou, Chi Chen, Ziyue Wang 외

Video temporal understanding is crucial for multimodal large language models (MLLMs) to reason over events in videos. Despite recent advances in general video understanding, current MLLMs still struggle with fine-grained…

Reinforcement Learning (RL)Video Understanding

VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding

2025-07-17 · Shihao Wang, Guo Chen, De-An Huang, Zhiqi Li 외

Recent studies have revealed that selecting informative and relevant video frames can significantly improve the performance of Video Large Language Models (Video-LLMs). Current methods, such as reducing inter-frame redun…

Video GroundingVideo Understanding

Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level

2024-11-15 · CVPR 2025 1 · Andong Deng, Tongjia Chen, Shoubin Yu, Taojiannan Yang 외

In this paper, we introduce Motion-Grounded Video Reasoning, a new motion understanding task that requires generating visual answers (video segmentation masks) according to the input question, and hence needs implicit sp…

BenchmarkingcounterfactualDescriptiveMultimodal Reasoning+4