paper-with-me

홈 › Papers

Temporal Concentration from Rollout Errors: Implicit Preference Optimization for Text-to-Video Diffusion

2026-07-30 · Henglin Liu, Fangyuan Kong, Jing Wang, Yizhou Lin, Nisha Huang, Chang Liu, Xintao Wang, Pengfei Wan, Kun Gai, Xiu Li arxiv

Recent advances in preference alignment for diffusion-based video generation, particularly via Direct Preference Optimization (DPO), have significantly improved visual quality. However, temporally sparse artifacts such as motion collapse, object flickering, and color oversaturation remain a major barrier to perceptual realism. Existing methods struggle with these issues due to two key limitations: (1) the preference attribution bottleneck, where offline human annotations are costly and fail to accurately capture learning dynamics, while online reward signals are rollout-aware but often unstable and biased; and (2) temporal credit misallocation, where uniformly applied supervision cannot effectively target the brief segments in which artifacts occur. To address these challenges, we propose concentrated Implicit Preference Optimization (cIPO), a post-training framework for video diffusion models. cIPO derives implicit preference signals directly from the denoising process: given a real video, the model adds forward noise and reconstructs it via iterative denoising, treating the original as the preferred sample and the reconstruction as the dispreferred one. This formulation captures inference-time errors without requiring human annotations or external reward models. Moreover, frame-level discrepancies between original and reconstructed videos reveal when failures occur. cIPO leverages this by computing temporal reconstruction errors and concentrating optimization on high-error segments, enabling more precise correction of failure-prone regions. Extensive experiments demonstrate that cIPO consistently enhances video authenticity and temporal coherence across multiple datasets, highlighting the effectiveness and efficiency of implicit preference with temporally concentrated optimization.

📄 PDF Abstract BibTeX arXiv:2607.28058

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

Thermalizer: Stable autoregressive neural emulation of spatiotemporal chaos

2025-03-24 · Chris Pedersen, Laure Zanna, Joan Bruna

Autoregressive surrogate models (or \textit{emulators}) of spatiotemporal systems provide an avenue for fast, approximate predictions, with broad applications across science and engineering. At inference time, however, t…

Denoising

Predictive Preference Learning from Human Interventions

2025-10-02 · Haoyuan Cai, Zhenghao Peng, Bolei Zhou arxiv

Learning from human involvement aims to incorporate the human subject to monitor and correct agent behavior errors. Although most interactive imitation learning methods focus on correcting the agent's action at the curre…

Autonomous Driving

GRAPE: Generalizing Robot Policy via Preference Alignment

2024-11-28 · Zijian Zhang, Kaiyuan Zheng, Zhaorun Chen, Joel Jang 외

Despite the recent advancements of vision-language-action (VLA) models on a variety of robotics tasks, they suffer from critical issues such as poor generalizability to unseen tasks, due to their reliance on behavior clo…

Vision-Language-Action

AMIR-GRPO: Inducing Implicit Preference Signals into GRPO

2026-01-07 · Amir Hossein Yari, Fajri Koto arxiv

Reinforcement learning has become the primary paradigm for aligning large language models (LLMs) on complex reasoning tasks, with group relative policy optimization (GRPO) widely used in large-scale post-training. Howeve…

Reinforcement LearningMathematical Reasoning

Pave-GRPO: Beyond Instantaneous Guidance through Principled Average Velocity Decomposition

2026-06-01 · Pengyang Ling, Jiazi Bu, Yujie Zhou, Yibin Wang 외 arxiv

Post-training via Group Relative Policy Optimization (GRPO) has emerged as a powerful paradigm for aligning flow-based generative models with human preferences. However, the iterative denoising nature of flow models incu…