paper-with-me

홈 › Papers

InfinityStory: Unlimited Video Generation with World Consistency and Character-Aware Shot Transitions

2026-03-04 · Mohamed Elmoghany, Liangbing Zhao, Xiaoqian Shen, Subhojyoti Mukherjee, Yang Zhou, Gang Wu, Viet Dac Lai, Seunghyun Yoon, Ryan Rossi, Abdullah Rashwan, Puneet Mathur, Varun Manjunatha, Daksh Dangi, Chien Nguyen, Nedim Lipka, Trung Bui, Krishna Kumar Singh, Ruiyi Zhang, Xiaolei Huang, Jaemin Cho, Yu Wang, Namyong Park, Zhengzhong Tu, Hongjie Chen, Hoda Eldardiry, Nesreen Ahmed, Thien Nguyen, Dinesh Manocha, Mohamed Elhoseiny, Franck Dernoncourt arxiv

Generating long-form storytelling videos with consistent visual narratives remains a significant challenge in video synthesis. We present a novel framework, dataset, and a model that address three critical limitations: background consistency across shots, seamless multi-subject shot-to-shot transitions, and scalability to hour-long narratives. Our approach introduces a background-consistent generation pipeline that maintains visual coherence across scenes while preserving character identity and spatial relationships. We further propose a transition-aware video synthesis module that generates smooth shot transitions for complex scenarios involving multiple subjects entering or exiting frames, going beyond the single-subject limitations of prior work. To support this, we contribute with a synthetic dataset of 10,000 multi-subject transition sequences covering underrepresented dynamic scene compositions. On VBench, InfinityStory achieves the highest Background Consistency (88.94), highest Subject Consistency (82.11), and the best overall average rank (2.80), showing improved stability, smoother transitions, and better temporal coherence.

📄 PDF Abstract BibTeX arXiv:2603.03646

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

Position: Interactive Generative Video as Next-Generation Game Engine

2025-03-21 · Jiwen Yu, Yiran Qin, Haoxuan Che, Quande Liu 외

Modern game development faces significant challenges in creativity and cost due to predetermined content in traditional game engines. Recent breakthroughs in video generation models, capable of synthesizing realistic and…

PositionVideo Generation

STAGE: A Stream-Centric Generative World Model for Long-Horizon Driving-Scene Simulation

2025-06-16 · Jiamin Wang, Yichen Yao, Xiang Feng, Hang Wu 외

The generation of temporally consistent, high-fidelity driving videos over extended horizons presents a fundamental challenge in autonomous driving world modeling. Existing approaches often suffer from error accumulation…

Autonomous DrivingDenoisingVideo Generation

Panacea: Panoramic and Controllable Video Generation for Autonomous Driving

2023-11-28 · CVPR 2024 1 · Yuqing Wen, Yucheng Zhao, Yingfei Liu, Fan Jia 외

The field of autonomous driving increasingly demands high-quality annotated training data. In this paper, we propose Panacea, an innovative approach to generate panoramic and controllable videos in driving scenarios, cap…

Autonomous DrivingVideo Generation

Geometric Reciprocity: Unlocking Self-Supervision for Stereoscopic Video Generation

2026-07-06 · Jingyi Lu, Kai Han arxiv

Monocular-to-stereo conversion synthesizes stereoscopic content from 2D videos for immersive 3D experiences. In modern Depth-Image-Based Rendering (DIBR) approaches, stereo inpainting of disocclusions is the critical bot…

Self-Supervised LearningVideo Generation

OmniDataComposer: A Unified Data Structure for Multimodal Data Fusion and Infinite Data Generation

2023-08-08 · Dongyang Yu, Shihao Wang, Yuan Fang, Wangpeng An

This paper presents OmniDataComposer, an innovative approach for multimodal data fusion and unlimited data generation with an intent to refine and uncomplicate interplay among diverse data modalities. Coming to the core …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Object TrackingOptical Character Recognition+6