paper-with-me

홈 › Papers

DeepMDP: Learning Continuous Latent Space Models for Representation Learning

2019-06-06 · Carles Gelada, Saurabh Kumar, Jacob Buckman, Ofir Nachum, Marc G. Bellemare

Many reinforcement learning (RL) tasks provide the agent with high-dimensional observations that can be simplified into low-dimensional continuous states. To formalize this process, we introduce the concept of a DeepMDP, a parameterized latent space model that is trained via the minimization of two tractable losses: prediction of rewards and prediction of the distribution over next latent states. We show that the optimization of these objectives guarantees (1) the quality of the latent space as a representation of the state space and (2) the quality of the DeepMDP as a model of the environment. We connect these results to prior work in the bisimulation literature, and explore the use of a variety of metrics. Our theoretical findings are substantiated by the experimental result that a trained DeepMDP recovers the latent structure underlying high-dimensional observations on a synthetic environment. Finally, we show that learning a DeepMDP as an auxiliary task in the Atari 2600 domain leads to large performance improvements over model-free RL.

📄 PDF Abstract BibTeX arXiv:1906.02736

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningReinforcement Learning (RL)Representation Learning

Similar Papers 제목 키워드 기반

Distillation of RL Policies with Formal Guarantees via Variational Abstraction of Markov Decision Processes (Technical Report)

2021-12-17 · Florent Delgrange, Ann Nowé, Guillermo A. Pérez

We consider the challenge of policy simplification and verification in the context of policies learned through reinforcement learning (RL) in continuous environments. In well-behaved settings, RL algorithms have converge…

Reinforcement Learning (RL)

CR-LSO: Convex Neural Architecture Optimization in the Latent Space of Graph Variational Autoencoder with Input Convex Neural Networks

2022-11-11 · Xuan Rao, Bo Zhao, Xiaosong Yi, Derong Liu

In neural architecture search (NAS) methods based on latent space optimization (LSO), a deep generative model is trained to embed discrete neural architectures into a continuous latent space. In this case, different opti…

Neural Architecture Search

Deep SPI: Safe Policy Improvement via World Models

2025-10-14 · Florent Delgrange, Raphael Avalos, Willem Röpke arxiv

Safe policy improvement (SPI) offers theoretical control over policy updates, yet existing guarantees largely concern offline, tabular reinforcement learning (RL). We study SPI in general online settings, when combined w…

Representation LearningReinforcement LearningOffline RL

Soft-IntroVAE for Continuous Latent space Image Super-Resolution

2023-07-18 · Zhi-Song Liu, Zijia Wang, Zhen Jia

Continuous image super-resolution (SR) recently receives a lot of attention from researchers, for its practical and flexible image scaling for various displays. Local implicit image representation is one of the methods t…

DenoisingImage RestorationImage Super-ResolutionSuper-Resolution

ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow

2026-07-30 · Dongxiu Liu, Haoyi Niu, Peng Cheng, Yuan Gao 외 arxiv

In the physical world we inhabit, space and time are fundamentally continuous. However, existing machine learning paradigms for world modeling are largely confined to discrete-time prediction, thereby exhibiting signific…

Image ReconstructionVideo Generation