paper-with-me

홈 › Papers

A Systematic Post-Train Framework for Video Generation

2026-04-28 · Zeyue Xue, Siming Fu, Jie Huang, Shuai Lu, Haoran Li, Yijun Liu, Yuming Li, Xiaoxuan He, Mengzhao Chen, Haoyang Huang, Nan Duan, Ping Luo arxiv

While large-scale video diffusion models have demonstrated impressive capabilities in generating high-resolution and semantically rich content, a significant gap remains between their pretraining performance and real-world deployment requirements due to critical issues such as prompt sensitivity, temporal inconsistency, and prohibitive inference costs. To bridge this gap, we propose a comprehensive post-training framework that systematically aligns pretrained models with user intentions through four synergistic stages: we first employ Supervised Fine-Tuning (SFT) to transform the base model into a stable instruction-following policy, followed by a Reinforcement Learning from Human Feedback (RLHF) stage that utilizes a novel Group Relative Policy Optimization (GRPO) method tailored for video diffusion to enhance perceptual quality and temporal coherence; subsequently, we integrate Prompt Enhancement via a specialized language model to refine user inputs, and finally address system efficiency through Inference Optimization. Together, these components provide a systematic approach to improving visual quality, temporal coherence, and instruction following, while preserving the controllability learned during pretraining. The result is a practical blueprint for building scalable post-training pipelines that are stable, adaptable, and effective in real-world deployment. Extensive experiments demonstrate that this unified pipeline effectively mitigates common artifacts and significantly improves controllability and visual aesthetics while adhering to strict sampling cost constraints.

📄 PDF Abstract BibTeX arXiv:2604.25427

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningInstruction FollowingVideo Generation

Similar Papers 제목 키워드 기반

TeleBoost: A Systematic Alignment Framework for High-Fidelity, Controllable, and Robust Video Generation

2026-02-07 · Yuanzhi Liang, Xuan'er Wu, Yirui Liu, Yijie Fang 외 arxiv

Post-training is the decisive step for converting a pretrained video generator into a production-oriented model that is instruction-following, controllable, and robust over long temporal horizons. This report presents a …

Reinforcement LearningVideo Generation

SoliReward: Mitigating Susceptibility to Reward Hacking and Annotation Noise in Video Generation Reward Models

2025-12-17 · Jiesong Lian, Ruizhe Zhong, Zixiang Zhou, Xiaoyue Mi 외 arxiv

Post-training alignment of video generation models with human preferences is a critical goal. Developing effective Reward Models (RMs) for this process faces significant methodological hurdles. Current data collection pa…

Video Generation

ASurvey: Spatiotemporal Consistency in Video Generation

2025-02-25 · Zhiyu Yin, Kehai Chen, Xuefeng Bai, Ruili Jiang 외

Video generation, by leveraging a dynamic visual generation method, pushes the boundaries of Artificial Intelligence Generated Content (AIGC). Video generation presents unique challenges beyond static image generation, r…

Image GenerationVideo Generation

DreaMoving: A Human Video Generation Framework based on Diffusion Models

2023-12-08 · Mengyang Feng, Jinlin Liu, Kai Yu, Yuan YAO 외

In this paper, we present DreaMoving, a diffusion-based controllable video generation framework to produce high-quality customized human videos. Specifically, given target identity and posture sequences, DreaMoving can g…

Video Generation

Learning Transferable Temporal Primitives for Video Reasoning via Synthetic Videos

2026-03-18 · Songtao Jiang, Sibo Song, Chenyi Zhou, Yuan Wang 외 arxiv

The transition from image to video understanding requires vision-language models (VLMs) to shift from recognizing static patterns to reasoning over temporal dynamics such as motion trajectories, speed changes, and state …

Video Generation