paper-with-me

홈 › Papers

Generalized Decision Transformer for Offline Hindsight Information Matching

2021-11-19 · Hiroki Furuta, Yutaka Matsuo, Shixiang Shane Gu

How to extract as much learning signal from each trajectory data has been a key problem in reinforcement learning (RL), where sample inefficiency has posed serious challenges for practical applications. Recent works have shown that using expressive policy function approximators and conditioning on future trajectory information -- such as future states in hindsight experience replay or returns-to-go in Decision Transformer (DT) -- enables efficient learning of multi-task policies, where at times online RL is fully replaced by offline behavioral cloning, e.g. sequence modeling. We demonstrate that all these approaches are doing hindsight information matching (HIM) -- training policies that can output the rest of trajectory that matches some statistics of future state information. We present Generalized Decision Transformer (GDT) for solving any HIM problem, and show how different choices for the feature function and the anti-causal aggregator not only recover DT as a special case, but also lead to novel Categorical DT (CDT) and Bi-directional DT (BDT) for matching different statistics of the future. For evaluating CDT and BDT, we define offline multi-task state-marginal matching (SMM) and imitation learning (IL) as two generic HIM problems, propose a Wasserstein distance loss as a metric for both, and empirically study them on MuJoCo continuous control benchmarks. CDT, which simply replaces anti-causal summation with anti-causal binning in DT, enables the first effective offline multi-task SMM algorithm that generalizes well to unseen and even synthetic multi-modal state-feature distributions. BDT, which uses an anti-causal second transformer as the aggregator, can learn to model any statistics of the future and outperforms DT variants in offline multi-task IL. Our generalized formulations from HIM and GDT greatly expand the role of powerful sequence modeling architectures in modern RL.

📄 PDF Abstract BibTeX arXiv:2111.10364

Code (1)

frt03/generalized_dt 공식 구현 pytorch

Tasks

continuous-controlContinuous ControlImitation LearningMuJoCoReinforcement Learning (RL)

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Adam 설명 없음
Position-Wise Feed-Forward Layer 설명 없음
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Residual Connection 설명 없음

Similar Papers 제목 키워드 기반

Skill Decision Transformer

2023-01-31 · Shyam Sudhakaran, Sebastian Risi

Recent work has shown that Large Language Models (LLMs) can be incredibly effective for offline reinforcement learning (RL) by representing the traditional RL problem as a sequence modelling problem (Chen et al., 2021; J…

D4RLDescriptiveOffline RLReinforcement Learning (RL)

Distributional Decision Transformer for Hindsight Information Matching

2021-09-29 · ICLR 2022 4 · Hiroki Furuta, Yutaka Matsuo, Shixiang Shane Gu

How to extract as much learning signal from each trajectory data has been a key problem in reinforcement learning (RL), where sample inefficiency has posed serious challenges for practical applications. Recent works have…

continuous-controlContinuous ControlImitation LearningMuJoCo+2

CEIL: Generalized Contextual Imitation Learning

2023-06-26 · NeurIPS 2023 11

In this paper, we present \textbf{C}ont\textbf{E}xtual \textbf{I}mitation \textbf{L}earning~(CEIL), a general and broadly applicable algorithm for imitation learning (IL). Inspired by the formulation of hindsight informa…

D4RLImitation LearningMuJoCo

Prompting Decision Transformers for Zero-Shot Reach-Avoid Policies

2025-05-25 · Kevin Li, Marinka Zitnik

Offline goal-conditioned reinforcement learning methods have shown promise for reach-avoid tasks, where an agent must reach a target state while avoiding undesirable regions of the state space. Existing approaches typica…

Dynamical-VAE-based Hindsight to Learn the Causal Dynamics of Factored-POMDPs

2024-11-12 · Chao Han, Debabrota Basu, Michael Mangan, Eleni Vasilaki 외

Learning representations of underlying environmental dynamics from partial observations is a critical challenge in machine learning. In the context of Partially Observable Markov Decision Processes (POMDPs), state repres…