paper-with-me

홈 › Papers

MoGAN: Improving Motion Quality in Video Diffusion via Few-Step Motion Adversarial Post-Training

2025-11-26 · Haotian Xue, Qi Chen, Zhonghao Wang, Xun Huang, Eli Shechtman, Jinrong Xie, Yongxin Chen arxiv

Video diffusion models achieve strong frame-level fidelity but still struggle with motion coherence, dynamics and realism, often producing jitter, ghosting, or implausible dynamics. A key limitation is that the standard denoising MSE objective provides no direct supervision on temporal consistency, allowing models to achieve low loss while still generating poor motion. We propose MoGAN, a motion-centric post-training framework that improves motion realism without reward models or human preference data. Built atop a 3-step distilled video diffusion model, we train a DiT-based optical-flow discriminator to differentiate real from generated motion, combined with a distribution-matching regularizer to preserve visual fidelity. With experiments on Wan2.1-T2V-1.3B, MoGAN substantially improves motion quality across benchmarks. On VBench, MoGAN boosts motion score by +7.3% over the 50-step teacher and +13.3% over the 3-step DMD model. On VideoJAM-Bench, MoGAN improves motion score by +7.4% over the teacher and +8.8% over DMD, while maintaining comparable or even better aesthetic and image-quality scores. A human study further confirms that MoGAN is preferred for motion quality (52% vs. 38% for the teacher; 56% vs. 29% for DMD). Overall, MoGAN delivers significantly more realistic motion without sacrificing visual fidelity or efficiency, offering a practical path toward fast, high-quality video generation. Project webpage is: https://xavihart.github.io/mogan.

📄 PDF Abstract BibTeX arXiv:2511.21592

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

Disentangling Content and Motion for Text-Based Neural Video Manipulation

2022-11-05 · Levent Karacan, Tolga Kerimoğlu, İsmail İnan, Tolga Birdal 외

Giving machines the ability to imagine possible new objects or scenes from linguistic descriptions and produce their realistic renderings is arguably one of the most challenging problems in computer vision. Recent advanc…

Motion Consistency Model: Accelerating Video Diffusion with Disentangled Motion-Appearance Distillation

2024-06-11 · Yuanhao Zhai, Kevin Lin, Zhengyuan Yang, Linjie Li 외

Image diffusion distillation achieves high-fidelity generation with very few sampling steps. However, applying these techniques directly to video diffusion often results in unsatisfactory frame quality due to the limited…

Denoising Reuse: Exploiting Inter-frame Motion Consistency for Efficient Video Latent Generation

2024-09-19 · Chenyu Wang, Shuo Yan, Yixuan Chen, Yujiang Wang 외

Video generation using diffusion-based models is constrained by high computational costs due to the frame-wise iterative diffusion process. This work presents a Diffusion Reuse MOtion (Dr. Mo) network to accelerate laten…

DenoisingVideo Generation

Real-Time Motion-Controllable Autoregressive Video Diffusion

2025-10-09 · Kesen Zhao, Jiaxin Shi, Beier Zhu, Junbao Zhou 외 arxiv

Real-time motion-controllable video generation remains challenging due to the inherent latency of bidirectional diffusion models and the lack of effective autoregressive (AR) approaches. Existing AR video diffusion model…

Text-to-Video GenerationReinforcement Learning

MogaNet: Multi-order Gated Aggregation Network

2022-11-07 · Siyuan Li, Zedong Wang, Zicheng Liu, Cheng Tan 외

By contextualizing the kernel as global as possible, Modern ConvNets have shown great potential in computer vision tasks. However, recent progress on \textit{multi-order game-theoretic interaction} within deep neural net…

3D Human Pose EstimationImage ClassificationInstance Segmentationobject-detection+5