paper-with-me

홈 › Papers

EasyAnimate: A High-Performance Long Video Generation Method based on Transformer Architecture

2024-05-29 · Jiaqi Xu, Xinyi Zou, Kunzhe Huang, Yunkuo Chen, Bo Liu, Mengli Cheng, Xing Shi, Jun Huang

This paper presents EasyAnimate, an advanced method for video generation that leverages the power of transformer architecture for high-performance outcomes. We have expanded the DiT framework originally designed for 2D image synthesis to accommodate the complexities of 3D video generation by incorporating a motion module block. It is used to capture temporal dynamics, thereby ensuring the production of consistent frames and seamless motion transitions. The motion module can be adapted to various DiT baseline methods to generate video with different styles. It can also generate videos with different frame rates and resolutions during both training and inference phases, suitable for both images and videos. Moreover, we introduce slice VAE, a novel approach to condense the temporal axis, facilitating the generation of long duration videos. Currently, EasyAnimate exhibits the proficiency to generate videos with 144 frames. We provide a holistic ecosystem for video production based on DiT, encompassing aspects such as data pre-processing, VAE training, DiT models training (both the baseline model and LoRA model), and end-to-end video inference. Code is available at: https://github.com/aigc-apps/EasyAnimate. We are continuously working to enhance the performance of our method.

📄 PDF Abstract BibTeX arXiv:2405.18991

Code (1)

aigc-apps/easyanimate 공식 구현 pytorch

Tasks

Image GenerationVideo Generation

Similar Papers 제목 키워드 기반

UniVid: The Open-Source Unified Video Model

2025-09-29 · Jiabin Luo, Junhui Lin, Zeyu Zhang, Biao Wu 외 arxiv

Unified video modeling that combines generation and understanding capabilities is increasingly important but faces two key challenges: maintaining semantic faithfulness during flow-based generation due to text-visual tok…

SWIFT: Sliding Window Reconstruction for Few-Shot Training-Free Generated Video Attribution

2026-03-09 · Chao Wang, Zijin Yang, Yaofei Wang, Yuang Qi 외 arxiv

Recent advancements in video generation technologies have been significant, resulting in their widespread application across multiple domains. However, concerns have been mounting over the potential misuse of generated c…

Video Generation

LongCat-Video Technical Report

2025-10-25 · Meituan LongCat Team, Xunliang Cai, Qilong Huang, Zhuoliang Kang 외 arxiv

Video generation is a critical pathway toward world models, with efficient long video inference as a key capability. Toward this end, we introduce LongCat-Video, a foundational video generation model with 13.6B parameter…

Video Generation

LongDiff: Training-Free Long Video Generation in One Go

2025-03-23 · CVPR 2025 1 · Zhuoling Li, Hossein Rahmani, Qiuhong Ke, Jun Liu

Video diffusion models have recently achieved remarkable results in video generation. Despite their encouraging performance, most of these models are mainly designed and trained for short video generation, leading to cha…

PositionVideo Generation

LVD-2M: A Long-take Video Dataset with Temporally Dense Captions

2024-10-14 · Tianwei Xiong, Yuqing Wang, Daquan Zhou, Zhijie Lin 외

The efficacy of video generation models heavily depends on the quality of their training datasets. Most previous video generation models are trained on short video clips, while recently there has been increasing interest…

Video CaptioningVideo Generation