paper-with-me

홈 › Papers

Skill Decision Transformer

2023-01-31 · Shyam Sudhakaran, Sebastian Risi

Recent work has shown that Large Language Models (LLMs) can be incredibly effective for offline reinforcement learning (RL) by representing the traditional RL problem as a sequence modelling problem (Chen et al., 2021; Janner et al., 2021). However many of these methods only optimize for high returns, and may not extract much information from a diverse dataset of trajectories. Generalized Decision Transformers (GDTs) (Furuta et al., 2021) have shown that utilizing future trajectory information, in the form of information statistics, can help extract more information from offline trajectory data. Building upon this, we propose Skill Decision Transformer (Skill DT). Skill DT draws inspiration from hindsight relabelling (Andrychowicz et al., 2017) and skill discovery methods to discover a diverse set of primitive behaviors, or skills. We show that Skill DT can not only perform offline state-marginal matching (SMM), but can discovery descriptive behaviors that can be easily sampled. Furthermore, we show that through purely reward-free optimization, Skill DT is still competitive with supervised offline RL approaches on the D4RL benchmark. The code and videos can be found on our project page: https://github.com/shyamsn97/skill-dt

📄 PDF Abstract BibTeX arXiv:2301.13573

Code (0)

등록된 구현이 없습니다.

Tasks

D4RLDescriptiveOffline RLReinforcement Learning (RL)

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Adam 설명 없음
Multi-Head Attention 설명 없음
Residual Connection 설명 없음

Similar Papers 제목 키워드 기반

LISA: Learning Interpretable Skill Abstractions from Language

2022-02-28 · Divyansh Garg, Skanda Vaidyanath, Kuno Kim, Jiaming Song 외

Learning policies that effectively utilize language instructions in complex, multi-task environments is an important problem in sequential decision-making. While it is possible to condition on the entire language instruc…

Decision MakingImitation LearningQuantizationSequential Decision Making

Unraveling the ARC Puzzle: Mimicking Human Solutions with Object-Centric Decision Transformer

2023-06-14 · JaeHyun Park, Jaegyun Im, Sanha Hwang, Mintaek Lim 외

In the pursuit of artificial general intelligence (AGI), we tackle Abstraction and Reasoning Corpus (ARC) tasks using a novel two-pronged approach. We employ the Decision Transformer in an imitation learning paradigm to …

ARCClusteringImitation Learningobject-detection+1

Learning to Plan & Schedule with Reinforcement-Learned Bimanual Robot Skills

2025-10-29 · Weikang Wan, Fabio Ramos, Xuning Yang, Caelan Garrett arxiv

Long-horizon contact-rich bimanual manipulation presents a significant challenge, requiring complex coordination involving a mixture of parallel execution and sequential collaboration between arms. In this paper, we intr…

Reinforcement Learning

Finding Skill Neurons in Pre-trained Transformer-based Language Models

2022-11-14 · Xiaozhi Wang, Kaiyue Wen, Zhengyan Zhang, Lei Hou 외

Transformer-based pre-trained language models have demonstrated superior performance on various natural language processing tasks. However, it remains unclear how the skills required to handle these tasks distribute amon…

Network Pruning

Skill Transformer: A Monolithic Policy for Mobile Manipulation

2023-08-19 · ICCV 2023 1 · Xiaoyu Huang, Dhruv Batra, Akshara Rai, Andrew Szot

We present Skill Transformer, an approach for solving long-horizon robotic tasks by combining conditional sequence modeling and skill modularity. Conditioned on egocentric and proprioceptive observations of a robot, Skil…

Task Planning