Video Generation
16개 벤치마크 · 논문 2,911편 · 이 태스크의 논문 보기 →
Benchmarks
UCF-101
BAIR Robot Pushing
Sky Time-lapse
LAION-400M
Taichi
How2Sign
Kinetics-700
MSR-VTT
TrailerFaces
YouTube Driving
Most implemented
GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
Consistency Models
Everybody Dance Now
Learning Temporal Coherence via Self-Supervision for GAN-based Video Generation
Video Diffusion Models
MoCoGAN: Decomposing Motion and Content for Video Generation
Papers
Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation
We present Vidu S2, which comprises Vidu S2-Avatar, a real-time interactive digital-character model, and Vidu S2-Editing, a real-time video editing model. Moreover, we explore the feasibility of real-time spatial video g…
Instruction FollowingVideo GenerationMask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout
Autoregressive (AR) video diffusion models have shown great potential in real-time video generation. Recent methods distill pretrained bidirectional video diffusion models into causal AR students through Distribution Mat…
Video GenerationRoLA: Rotary-Positioned Low-Rank Linear Attention for Efficient Diffusion Transformers
Diffusion Transformers (DiTs) achieve strong video generation quality, but their dense spatiotemporal self-attention scales quadratically with sequence length and quickly becomes the dominant inference bottleneck. Sparse…
Video GenerationMulti-Grid Post-Training for Long-Form Multi-Shot Video Generation
Generating long-form multi-shot videos requires coherent within-shot motion and visually consistent narratives across shots. Existing video generators favor continuous motion and struggle to present complete shot sets wh…
Video GenerationTourPhysics: Bringing Physics to World Models for Exploration and Manipulation from a Single Image
Interactive visual world models must distinguish observation from physical intervention. Camera motion reveals new surfaces, whereas intervention changes object motion, contact, and deformation. Current video world model…
Video GenerationPRISM-Bench: An Audio-Centric Diagnostic Benchmark for Text-to-Audio-Video Generation
Text-to-audio-video (T2AV) generation has advanced rapidly, but its evaluation still underestimates the audio modality. Existing benchmarks either treat audio as an auxiliary component of video quality or assess it in is…
Audio GenerationVideo Generation