paper-with-me

홈 › Papers

LongVie 2: Multimodal Controllable Ultra-Long Video World Model

2025-12-15 · Jianxiong Gao, Zhaoxi Chen, Xian Liu, Junhao Zhuang, Chengming Xu, Jianfeng Feng, Yu Qiao, Yanwei Fu, Chenyang Si, Ziwei Liu arxiv

Building video world models upon pretrained video generation systems represents an important yet challenging step toward general spatiotemporal intelligence. A world model should possess three essential properties: controllability, long-term visual quality, and temporal consistency. To this end, we take a progressive approach-first enhancing controllability and then extending toward long-term, high-quality generation. We present LongVie 2, an end-to-end autoregressive framework trained in three stages: (1) Multi-modal guidance, which integrates dense and sparse control signals to provide implicit world-level supervision and improve controllability; (2) Degradation-aware training on the input frame, bridging the gap between training and long-term inference to maintain high visual quality; and (3) History-context guidance, which aligns contextual information across adjacent clips to ensure temporal consistency. We further introduce LongVGenBench, a comprehensive benchmark comprising 100 high-resolution one-minute videos covering diverse real-world and synthetic environments. Extensive experiments demonstrate that LongVie 2 achieves state-of-the-art performance in long-range controllability, temporal coherence, and visual fidelity, and supports continuous video generation lasting up to five minutes, marking a significant step toward unified video world modeling.

📄 PDF Abstract BibTeX arXiv:2512.13604

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

LongVie: Multimodal-Guided Controllable Ultra-Long Video Generation

2025-08-05 · Jianxiong Gao, Zhaoxi Chen, Xian Liu, Jianfeng Feng 외 arxiv

Controllable ultra-long video generation is a fundamental yet challenging task. Although existing methods are effective for short clips, they struggle to scale due to issues such as temporal inconsistency and visual degr…

Video Generation

M3-CVC: Controllable Video Compression with Multimodal Generative Models

2024-11-24 · Rui Wan, Qi Zheng, Yibo Fan

Traditional and neural video codecs commonly encounter limitations in controllability and generality under ultra-low-bitrate coding scenarios. To overcome these challenges, we propose M3-CVC, a controllable video compres…

Video Compression

Bridging Modalities, Spanning Time: Structured Memory for Ultra-Long Agentic Video Reasoning

2026-05-08 · Jiazheng Li, Chi-Hao Wu, Yunze Liu, Kaize Ding 외 arxiv

Understanding ultra-long videos such as egocentric recordings, live streams, or surveillance footage spanning days to weeks, remains a challenge. For current multimodal LLMs: even with million-token context windows, fram…

Cross-Modal Retrieval

AutoCut: End-to-end advertisement video editing based on multimodal discretization and controllable generation

2026-03-30 · Milton Zhou, Sizhong Qin, Yongzhi Li, Quan Chen 외 arxiv

Short-form videos have become a primary medium for digital advertising, requiring scalable and efficient content creation. However, current workflows and AI tools remain disjoint and modality-specific, leading to high pr…

Video Object Segmentation-Aware Audio Generation

2025-09-30 · Ilpo Viertola, Vladimir Iashin, Esa Rahtu arxiv

Existing multimodal audio generation models often lack precise user control, which limits their applicability in professional Foley workflows. In particular, these models focus on the entire video and do not provide prec…

Video Object SegmentationAudio Generation