paper-with-me

홈 › Papers

TempSamp-R1: Effective Temporal Sampling with Reinforcement Fine-Tuning for Video LLMs

2025-09-22 · Yunheng Li, Jing Cheng, Shaoyong Jia, Hangyi Kuang, Shaohui Jiao, Qibin Hou, Ming-Ming Cheng arxiv

This paper introduces TempSamp-R1, a new reinforcement fine-tuning framework designed to improve the effectiveness of adapting multimodal large language models (MLLMs) to video temporal grounding tasks. We reveal that existing reinforcement learning methods, such as Group Relative Policy Optimization (GRPO), rely on on-policy sampling for policy updates. However, in tasks with large temporal search spaces, this strategy becomes both inefficient and limited in performance, as it often fails to identify temporally accurate solutions. To address this limitation, TempSamp-R1 leverages ground-truth annotations as off-policy supervision to provide temporally precise guidance, effectively compensating for the sparsity and misalignment in on-policy solutions. To further stabilize training and reduce variance in reward-based updates, TempSamp-R1 provides a non-linear soft advantage computation method that dynamically reshapes the reward feedback via an asymmetric transformation. By employing a hybrid Chain-of-Thought (CoT) training paradigm, TempSamp-R1 optimizes a single unified model to support both CoT and non-CoT inference modes, enabling efficient handling of queries with varying reasoning complexity. Experimental results demonstrate that TempSamp-R1 outperforms GRPO-based baselines, establishing new state-of-the-art performance on benchmark datasets: Charades-STA (R1@0.7: 52.9%, +2.7%), ActivityNet Captions (R1@0.5: 56.0%, +5.3%), and QVHighlights (mAP: 30.0%, +3.0%). Moreover, TempSamp-R1 shows robust few-shot generalization capabilities under limited data. Code: https://github.com/HVision-NKU/TempSamp-R1

📄 PDF Abstract BibTeX arXiv:2509.18056

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Temporal Sampling for Forgotten Reasoning in LLMs

2025-05-26 · Yuetai Li, Zhangchen Xu, Fengqing Jiang, Bhaskar Ramasubramanian 외

Fine-tuning large language models (LLMs) is intended to improve their reasoning capabilities, yet we uncover a counterintuitive effect: models often forget how to solve problems they previously answered correctly during …

FrameMind: Frame-Interleaved Video Reasoning via Reinforcement Learning

2025-09-28 · Haonan Ge, Yiwei Wang, Kai-Wei Chang, Hang Wu 외 arxiv

Current video understanding models rely on fixed frame sampling strategies, processing predetermined visual inputs regardless of the specific reasoning requirements of each question. This static approach limits their abi…

Reinforcement Learning

Learning Dense Reward with Temporal Variant Self-Supervision

2022-05-20 · Yuning Wu, Jieliang Luo, Hui Li

Rewards play an essential role in reinforcement learning. In contrast to rule-based game environments with well-defined reward functions, complex real-world robotic applications, such as contact-rich manipulation, lack e…

Contact-rich ManipulationSelf-Supervised Learning

Reinforcement Learning for Sampling on Temporal Medical Imaging Sequences

2023-08-28 · Zhishen Huang

Accelerated magnetic resonance imaging resorts to either Fourier-domain subsampling or better reconstruction algorithms to deal with fewer measurements while still generating medical images of high quality. Determining t…

Image ReconstructionQ-Learningreinforcement-learningReinforcement Learning+1

TSPO: Temporal Sampling Policy Optimization for Long-form Video Language Understanding

2025-08-06 · Canhui Tang, Zifan Han, Hongbo Sun, Sanping Zhou 외 arxiv

Multimodal Large Language Models (MLLMs) have demonstrated significant progress in vision-language tasks, yet they still face challenges when processing long-duration video inputs. The limitation arises from MLLMs' conte…

Reinforcement Learning