paper-with-me

Papers

StateSpaceDiffuser: Bringing Long Context to Diffusion World Models

2025-05-28 · Nedko Savov, Naser Kazemi, Deheng Zhang, Danda Pani Paudel, Xi Wang, Luc van Gool

World models have recently become promising tools for predicting realistic visuals based on actions in complex environments. However, their reliance on a short sequence of observations causes them to quickly lose track of context. As a result, visual consistency breaks down after just a few steps, and generated scenes no longer reflect information seen earlier. This limitation of the state-of-the-art diffusion-based world models comes from their lack of a lasting environment state. To address this problem, we introduce StateSpaceDiffuser, where a diffusion model is enabled to perform on long-context tasks by integrating a sequence representation from a state-space model (Mamba), representing the entire interaction history. This design restores long-term memory without sacrificing the high-fidelity synthesis of diffusion models. To rigorously measure temporal consistency, we develop an evaluation protocol that probes a model's ability to reinstantiate seen content in extended rollouts. Comprehensive experiments show that StateSpaceDiffuser significantly outperforms a strong diffusion-only baseline, maintaining a coherent visual context for an order of magnitude more steps. It delivers consistent views in both a 2D maze navigation and a complex 3D environment. These results establish that bringing state-space representations into diffusion models is highly effective in demonstrating both visual details and long-term memory.

📄 PDF Abstract BibTeX arXiv:2505.22246

Code (0)

등록된 구현이 없습니다.

Tasks

Mamba

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Jurassic World Remake: Bringing Ancient Fossils Back to Life via Zero-Shot Long Image-to-Image Translation

2023-08-14 · Alexander Martin, Haitian Zheng, Jie An, Jiebo Luo

With a strong understanding of the target domain from natural language, we produce promising results in translating across large domain gaps and bringing skeletons back to life. In this work, we use text-guided latent di…

Image-to-Image Translation

You Only Erase Once: Erasing Anything without Bringing Unexpected Content

2026-03-29 · Yixing Zhu, Qing Zhang, Wenju Xu, Wei-Shi Zheng arxiv

We present YOEO, an approach for object erasure. Unlike recent diffusion-based methods which struggle to erase target objects without generating unexpected content within the masked regions due to lack of sufficient pair…

Diffusion-Based Action Recognition Generalizes to Untrained Domains

2025-09-10 · Rogerio Guimaraes, Frank Xiao, Pietro Perona, Markus Marks arxiv

Humans can recognize the same actions despite large context and viewpoint variations, such as differences between species (walking in spiders vs. horses), viewpoints (egocentric vs. third-person), and contexts (real life…

Action Recognition

VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models

2024-03-10 · Wenhao Wang, Yi Yang

The arrival of Sora marks a new era for text-to-video diffusion models, bringing significant advancements in video generation and potential applications. However, Sora, along with other text-to-video diffusion models, is…

Copy DetectionImage GenerationPrompt EngineeringText-to-Video Generation+1

Iris: Bringing Real-World Priors into Diffusion Model for Monocular Depth Estimation

2026-03-17 · Xinhao Cai, Gensheng Pei, Zeren Sun, Yazhou Yao 외 arxiv

In this paper, we propose \textbf{Iris}, a deterministic framework for Monocular Depth Estimation (MDE) that integrates real-world priors into the diffusion model. Conventional feed-forward methods rely on massive traini…

Monocular Depth Estimation