paper-with-me

홈 › Papers

One-Minute Video Generation with Test-Time Training

2025-04-07 · CVPR 2025 1 · Karan Dalal, Daniel Koceja, Gashon Hussein, Jiarui Xu, Yue Zhao, Youjin Song, Shihao Han, Ka Chun Cheung, Jan Kautz, Carlos Guestrin, Tatsunori Hashimoto, Sanmi Koyejo, Yejin Choi, Yu Sun, Xiaolong Wang

Transformers today still struggle to generate one-minute videos because self-attention layers are inefficient for long context. Alternatives such as Mamba layers struggle with complex multi-scene stories because their hidden states are less expressive. We experiment with Test-Time Training (TTT) layers, whose hidden states themselves can be neural networks, therefore more expressive. Adding TTT layers into a pre-trained Transformer enables it to generate one-minute videos from text storyboards. For proof of concept, we curate a dataset based on Tom and Jerry cartoons. Compared to baselines such as Mamba~2, Gated DeltaNet, and sliding-window attention layers, TTT layers generate much more coherent videos that tell complex stories, leading by 34 Elo points in a human evaluation of 100 videos per method. Although promising, results still contain artifacts, likely due to the limited capability of the pre-trained 5B model. The efficiency of our implementation can also be improved. We have only experimented with one-minute videos due to resource constraints, but the approach can be extended to longer videos and more complex stories. Sample videos, code and annotations are available at: https://test-time-training.github.io/video-dit

📄 PDF Abstract BibTeX arXiv:2504.05298

Code (0)

등록된 구현이 없습니다.

Tasks

MambaVideo Generation

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Position-Wise Feed-Forward Layer 설명 없음
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Multi-Head Attention 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Adam 설명 없음

Similar Papers 제목 키워드 기반

Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models

2026-02-24 · Christian Simon, Masato Ishii, Wei-Yao Wang, Koichi Saito 외 arxiv

Scaling multimodal alignment between video and audio is challenging, particularly due to limited data and the mismatch between text descriptions and frame-level video information. In this work, we tackle the scaling chal…

Audio Generation

Magic 1-For-1: Generating One Minute Video Clips within One Minute

2025-02-11 · Hongwei Yi, Shitong Shao, Tian Ye, Jiantong Zhao 외

In this technical report, we present Magic 1-For-1 (Magic141), an efficient video generation model with optimized memory consumption and inference latency. The key idea is simple: factorize the text-to-video generation t…

Image GenerationImage to Video GenerationText to Image GenerationText-to-Image Generation+2

LongCat-Video Technical Report

2025-10-25 · Meituan LongCat Team, Xunliang Cai, Qilong Huang, Zhuoliang Kang 외 arxiv

Video generation is a critical pathway toward world models, with efficient long video inference as a key capability. Toward this end, we introduce LongCat-Video, a foundational video generation model with 13.6B parameter…

Video Generation

LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV

2026-05-25 · Tengfei Liu, Yang Shi, Xuanyu Zhu, Jiafu Tang 외 arxiv

Audio-visual generation is rapidly advancing from short clips to minute-long content, while existing evaluation protocols remain largely confined to short-form settings. Existing benchmarks primarily focus on 5--10 secon…

BIFE: Better Interaction, Fewer Errors for Minute-Long Video Generation

2025-11-28 · Zeyu Zhang, Jinyuan Mao, Shuning Chang, Yuanyu He 외 arxiv

Long video generation is a critical step toward building realistic world models, requiring both high visual fidelity and long-range interaction consistency. Recent autoregressive diffusion models enable long-horizon gene…

Video Generation