paper-with-me

Papers

Diffusion for World Modeling: Visual Details Matter in Atari

2024-05-20 · Eloi Alonso, Adam Jelley, Vincent Micheli, Anssi Kanervisto, Amos Storkey, Tim Pearce, François Fleuret

World models constitute a promising approach for training reinforcement learning agents in a safe and sample-efficient manner. Recent world models predominantly operate on sequences of discrete latent variables to model environment dynamics. However, this compression into a compact discrete representation may ignore visual details that are important for reinforcement learning. Concurrently, diffusion models have become a dominant approach for image generation, challenging well-established methods modeling discrete latents. Motivated by this paradigm shift, we introduce DIAMOND (DIffusion As a Model Of eNvironment Dreams), a reinforcement learning agent trained in a diffusion world model. We analyze the key design choices that are required to make diffusion suitable for world modeling, and demonstrate how improved visual details can lead to improved agent performance. DIAMOND achieves a mean human normalized score of 1.46 on the competitive Atari 100k benchmark; a new best for agents trained entirely within a world model. We further demonstrate that DIAMOND's diffusion world model can stand alone as an interactive neural game engine by training on static Counter-Strike: Global Offensive gameplay. To foster future research on diffusion for world modeling, we release our code, agents, videos and playable world models at https://diamond-wm.github.io.

📄 PDF Abstract BibTeX arXiv:2405.12399

Code (1)

eloialonso/diamond 공식 구현 pytorch

Tasks

Image Generationreinforcement-learningReinforcement Learning

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Video-GPT via Next Clip Diffusion

2025-05-18 · Shaobin Zhuang, Zhipeng Huang, Ying Zhang, Fangyikang Wang 외

GPT has shown its remarkable success in natural language processing. However, the language sequence is not sufficient to describe spatial-temporal details in the visual world. Alternatively, the video sequence is good at…

DenoisingImage AnimationPredictionVideo Classification+4

What Matters When Repurposing Diffusion Models for General Dense Perception Tasks?

2024-03-10 · Guangkai Xu, Yongtao Ge, MingYu Liu, Chengxiang Fan 외

Extensive pre-training with large data is indispensable for downstream geometry and semantic visual perception tasks. Thanks to large-scale text-to-image (T2I) pretraining, recent works show promising results by simply f…

Depth EstimationImage MattingImage SegmentationMonocular Depth Estimation+3

Prior Does Matter: Visual Navigation via Denoising Diffusion Bridge Models

2025-04-14 · CVPR 2025 1 · Hao Ren, Yiming Zeng, Zetong Bi, Zhaoliang Wan 외

Recent advancements in diffusion-based imitation learning, which show impressive performance in modeling multimodal distributions and training stability, have led to substantial progress in various robot learning tasks. …

Action GenerationDenoisingImitation LearningVisual Navigation

Distribution of Matter from Singularity with Spherical Symmetry Using Fick Diffusion

2006-05-12 · Ahmet Mecit Öztaş, Michael L. Smith

A model is presented allowing calculation of energy and matter distribution in the Universe after expansion from singularity without introduction of expansion energy. Beginning with Fick's law of diffusion, we solve the …

A Machine Learning Approach For Identifying Patients with Mild Traumatic Brain Injury Using Diffusion MRI Modeling

2017-08-27 · Shervin Minaee, Yao Wang, Sohae Chung, Xiuyuan Wang 외

While diffusion MRI has been extremely promising in the study of MTBI, identifying patients with recent MTBI remains a challenge. The literature is mixed with regard to localizing injury in these patients, however, gray …

BIG-bench Machine LearningDiffusion MRI