paper-with-me

홈 › Papers

GigaVideo-1: Advancing Video Generation via Automatic Feedback with 4 GPU-Hours Fine-Tuning

2025-06-12 · Xiaoyi Bao, Jindi Lv, XiaoFeng Wang, Zheng Zhu, Xinze Chen, Yukun Zhou, Jiancheng Lv, Xingang Wang, Guan Huang

Recent progress in diffusion models has greatly enhanced video generation quality, yet these models still require fine-tuning to improve specific dimensions like instance preservation, motion rationality, composition, and physical plausibility. Existing fine-tuning approaches often rely on human annotations and large-scale computational resources, limiting their practicality. In this work, we propose GigaVideo-1, an efficient fine-tuning framework that advances video generation without additional human supervision. Rather than injecting large volumes of high-quality data from external sources, GigaVideo-1 unlocks the latent potential of pre-trained video diffusion models through automatic feedback. Specifically, we focus on two key aspects of the fine-tuning process: data and optimization. To improve fine-tuning data, we design a prompt-driven data engine that constructs diverse, weakness-oriented training samples. On the optimization side, we introduce a reward-guided training strategy, which adaptively weights samples using feedback from pre-trained vision-language models with a realism constraint. We evaluate GigaVideo-1 on the VBench-2.0 benchmark using Wan2.1 as the baseline across 17 evaluation dimensions. Experiments show that GigaVideo-1 consistently improves performance on almost all the dimensions with an average gain of about 4% using only 4 GPU-hours. Requiring no manual annotations and minimal real data, GigaVideo-1 demonstrates both effectiveness and efficiency. Code, model, and data will be publicly available.

📄 PDF Abstract BibTeX arXiv:2506.10639

Code (0)

등록된 구현이 없습니다.

Tasks

GPUVideo Generation

Methods 이 논문이 사용한 방법론

Focus 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation

2024-06-21 · Xuan He, Dongfu Jiang, Ge Zhang, Max Ku 외

The recent years have witnessed great advances in video generation. However, the development of automatic video metrics is lagging significantly behind. None of the existing metric is able to provide reliable scores over…

Video GenerationVideo Quality Assessment

Evaluation Framework for Feedback Generation Methods in Skeletal Movement Assessment

2024-04-14 · Tal Hakim

The application of machine-learning solutions to movement assessment from skeleton videos has attracted significant research attention in recent years. This advancement has made rehabilitation at home more accessible, ut…

Bridging Creative Intent and Visual Quality: Creator-Driven Recurrent Video Generation with Agentic Feedback Loops

2026-06-17 · Denis Savytski, Aiden Lei, Heding Liu, Warren Yang 외 arxiv

Generative AI has made content creation increasingly accessible, but many AI-generated videos lack narrative coherence and creative direction, issues that become more substantial at longer durations. Unlike coding, where…

Video Generation

We'll Fix it in Post: Improving Text-to-Video Generation with Neuro-Symbolic Feedback

2025-04-24 · Minkyu Choi, S P Sharan, Harsh Goel, Sahil Shah 외

Current text-to-video (T2V) generation models are increasingly popular due to their ability to produce coherent videos from textual prompts. However, these models often struggle to generate semantically and temporally co…

Text-to-Video GenerationVideo Generation

RLGF: Reinforcement Learning with Geometric Feedback for Autonomous Driving Video Generation

2025-09-20 · Tianyi Yan, Wencheng Han, Xia Zhou, Xueyang Zhang 외 arxiv

Synthetic data is crucial for advancing autonomous driving (AD) systems, yet current state-of-the-art video generation models, despite their visual realism, suffer from subtle geometric distortions that limit their utili…

Reinforcement Learning3D Object DetectionAutonomous DrivingVideo Generation