paper-with-me

Papers

Modeling the Long Term Future in Model-Based Reinforcement Learning

2019-05-01 · ICLR 2019 5 · Nan Rosemary Ke, Amanpreet Singh, Ahmed Touati, Anirudh Goyal, Yoshua Bengio, Devi Parikh, Dhruv Batra

In model-based reinforcement learning, the agent interleaves between model learning and planning. These two components are inextricably intertwined. If the model is not able to provide sensible long-term prediction, the executed planer would exploit model flaws, which can yield catastrophic failures. This paper focuses on building a model that reasons about the long-term future and demonstrates how to use this for efficient planning and exploration. To this end, we build a latent-variable autoregressive model by leveraging recent ideas in variational inference. We argue that forcing latent variables to carry future information through an auxiliary task substantially improves long-term predictions. Moreover, by planning in the latent space, the planner's solution is ensured to be within regions where the model is valid. An exploration strategy can be devised by searching for unlikely trajectories under the model. Our methods achieves higher reward faster compared to baselines on a variety of tasks and environments in both the imitation learning and model-based reinforcement learning settings.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Imitation LearningModel-based Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)validVariational Inference

Similar Papers 제목 키워드 기반

Diffusion World Model: Future Modeling Beyond Step-by-Step Rollout for Offline Reinforcement Learning

2024-02-05 · Zihan Ding, Amy Zhang, Yuandong Tian, Qinqing Zheng

We introduce Diffusion World Model (DWM), a conditional diffusion model capable of predicting multistep future states and rewards concurrently. As opposed to traditional one-step dynamics models, DWM offers long-horizon …

D4RLQ-Learning

PRISM: Preference Refinement via Implicit Scene Modeling for 3D Vision-Language Preference-Based Reinforcement Learning

2025-03-13 · Yirong Sun, Yanjun Chen

We propose PRISM, a novel framework designed to overcome the limitations of 2D-based Preference-Based Reinforcement Learning (PBRL) by unifying 3D point cloud modeling and future-aware preference refinement. At its core,…

Autonomous NavigationDecision MakingLanguage ModelingLanguage Modelling+2

Imitation Learning for Human Pose Prediction

2019-09-08 · ICCV 2019 10 · Borui Wang, Ehsan Adeli, Hsu-kuang Chiu, De-An Huang 외

Modeling and prediction of human motion dynamics has long been a challenging problem in computer vision, and most existing methods rely on the end-to-end supervised training of various architectures of recurrent neural n…

Deep Reinforcement LearningHuman Pose ForecastingImitation LearningPose Prediction+4

Chat More If You Like: Dynamic Cue Words Planning to Flow Longer Conversations

2018-11-19 · Lili Yao, Ruijian Xu, Chao Li, Dongyan Zhao 외

To build an open-domain multi-turn conversation system is one of the most interesting and challenging tasks in Artificial Intelligence. Many research efforts have been dedicated to building such dialogue systems, yet few…

DiversityReinforcement Learning

Transformer-Enhanced Reinforcement Learning: Fundamentals and Applications in Communication Networks

2026-05-26 · Nguyen Cong Luong, Shaohan Feng, Nguyen Duc Hai, Zeping Sui 외 arxiv

Reinforcement Learning (RL) has long been a powerful solution to various problems in communication networks. However, traditional RL models still face with several limitations. Not only do they rely on large numbers of i…

Reinforcement LearningSemantic Communication