paper-with-me

홈 › Papers

Temporal Preference Optimization for Long-Form Video Understanding

2025-01-23 · Rui Li, Xiaohan Wang, Yuhui Zhang, Zeyu Wang, Serena Yeung-Levy

Despite significant advancements in video large multimodal models (video-LMMs), achieving effective temporal grounding in long-form videos remains a challenge for existing models. To address this limitation, we propose Temporal Preference Optimization (TPO), a novel post-training framework designed to enhance the temporal grounding capabilities of video-LMMs through preference learning. TPO adopts a self-training approach that enables models to differentiate between well-grounded and less accurate temporal responses by leveraging curated preference datasets at two granularities: localized temporal grounding, which focuses on specific video segments, and comprehensive temporal grounding, which captures extended temporal dependencies across entire video sequences. By optimizing on these preference datasets, TPO significantly enhances temporal understanding while reducing reliance on manually annotated data. Extensive experiments on three long-form video understanding benchmarks--LongVideoBench, MLVU, and Video-MME--demonstrate the effectiveness of TPO across two state-of-the-art video-LMMs. Notably, LLaVA-Video-TPO establishes itself as the leading 7B model on the Video-MME benchmark, underscoring the potential of TPO as a scalable and efficient solution for advancing temporal reasoning in long-form video understanding. Project page: https://ruili33.github.io/tpo_website.

📄 PDF Abstract BibTeX arXiv:2501.13919

Code (0)

등록된 구현이 없습니다.

Tasks

FormMMEVideo MMEVideo Understanding

Similar Papers 제목 키워드 기반

VideoPASTA: 7K Preference Pairs That Matter for Video-LLM Alignment

2025-04-18 · Yogesh Kulkarni, Pooyan Fazli

Video-language models (Video-LLMs) excel at understanding video content but struggle with spatial relationships, temporal ordering, and cross-frame continuity. To address these limitations, we introduce VideoPASTA (Prefe…

MVBench

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models

2025-04-17 · Haojian Huang, Haodong Chen, Shengqiong Wu, Meng Luo 외

Large Video Models (LVMs) built upon Large Language Models (LLMs) have shown promise in video understanding but often suffer from misalignment with human intuition and video hallucination issues. To address these challen…

HallucinationVideo Understanding

Self-alignment of Large Video Language Models with Refined Regularized Preference Optimization

2025-04-16 · Pritam Sarkar, Ali Etemad

Despite recent advances in Large Video Language Models (LVLMs), they still struggle with fine-grained temporal understanding, hallucinate, and often make simple mistakes on even simple video question-answering tasks, all…

HallucinationQuestion AnsweringVideo Question AnsweringVideo Understanding

TEMPLE:Temporal Preference Learning of Video LLMs via Difficulty Scheduling and Pre-SFT Alignment

2025-03-21 · Shicheng Li, Lei LI, Kun Ouyang, Shuhuai Ren 외

Video Large Language Models (Video LLMs) have achieved significant success by leveraging a two-stage paradigm: pretraining on large-scale video-text data for vision-language alignment, followed by supervised fine-tuning …

Scheduling

Temporal Concentration from Rollout Errors: Implicit Preference Optimization for Text-to-Video Diffusion

2026-07-30 · Henglin Liu, Fangyuan Kong, Jing Wang, Yizhou Lin 외 arxiv

Recent advances in preference alignment for diffusion-based video generation, particularly via Direct Preference Optimization (DPO), have significantly improved visual quality. However, temporally sparse artifacts such a…

Video Generation