paper-with-me

Papers

MosaicMem: Hybrid Spatial Memory for Controllable Video World Models

2026-03-17 · Wei Yu, Runjia Qian, Yumeng Li, Liquan Wang, Songheng Yin, Sri Siddarth Chakaravarthy P, Dennis Anthony, Yang Ye, Yidi Li, Weiwei Wan, Animesh Garg arxiv

Video diffusion models are moving beyond short, plausible clips toward world simulators that must remain consistent under camera motion, revisits, and intervention. Yet spatial memory remains a key bottleneck: explicit 3D structures can improve reprojection-based consistency but struggle to depict moving objects, while implicit memory often produces inaccurate camera motion even with correct poses. We propose Mosaic Memory (MosaicMem), a hybrid spatial memory that lifts patches into 3D for reliable localization and targeted retrieval, while exploiting the model's native conditioning to preserve prompt-following generation. MosaicMem composes spatially aligned patches in the queried view via a patch-and-compose interface, preserving what should persist while allowing the model to inpaint what should evolve. With PRoPE camera conditioning and two new memory alignment methods, experiments show improved pose adherence compared to implicit memory and stronger dynamic modeling than explicit baselines. MosaicMem further enables minute-level navigation, memory-based scene editing, and autoregressive rollout.

📄 PDF Abstract BibTeX arXiv:2603.17117

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Attend Locally, Remember Linearly: Linear Attention as Cross-Frame Memory for Autoregressive Video Diffusion

2026-05-15 · Kunyang Li, Mubarak Shah, Yuzhang Shang arxiv

Autoregressive (AR) video diffusion is a powerful paradigm for streaming and interactive video generation. However, its reliance on softmax self-attention leads to quadratic compute complexity in sequence length and memo…

Video Generation

Controllable Hybrid Captioner for Improved Long-form Video Understanding

2025-07-22 · Kuleen Sasse, Efsun Sarioglu Kayi, Arun Reddy arxiv

Video data, especially long-form video, is extremely dense and high-dimensional. Text-based summaries of video content offer a way to represent query-relevant content in a much more compact manner than raw video. In addi…

Natural Language Queries

Memorize When Needed: Decoupled Memory Control for Spatially Consistent Long-Horizon Video Generation

2026-04-20 · Yanjun Guo, Zhengqiang Zhang, Pengfei Wang, Xinyue Liang 외 arxiv

Spatially consistent long-horizon video generation aims to maintain temporal and spatial consistency along predefined camera trajectories. Existing methods mostly entangle memory modeling with video generation, leading t…

Video Generation

Temporal Hallucinating for Action Recognition With Few Still Images

2018-06-01 · CVPR 2018 6 · Yali Wang, Lei Zhou, Yu Qiao

Action recognition in still images has been recently promoted by deep learning. However, the success of these deep models heavily depends on huge amount of training images for various action categories, which may not be …

Action RecognitionAction Recognition In Still ImagesDomain AdaptationTemporal Action Localization

AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories

2026-02-16 · Zun Wang, Han Lin, Jaehong Yoon, Jaemin Cho 외 arxiv

Maintaining spatial world consistency over long horizons remains a central challenge for camera-controllable video generation. Existing memory-based approaches often condition generation on globally reconstructed 3D scen…

Depth EstimationVideo Generation