paper-with-me

홈 › Papers

Learning Hierarchical Skill Policies with Offline Quality-Diversity Reinforcement Learning

2026-08-20 · Tanachai Anakewat, Takayuki Osa, Tatsuya Harada arxiv

Recent studies investigate how to leverage pre-collected datasets to improve the policy performance and sample efficiency of RL. One promising approach to achieve this goal is to employ a two-stage strategy: In the first stage, diverse skills are extracted as a low-level policy from a given dataset, and a high-level policy is trained to solve a specific task in the second stage. Typically, extraction of the low-level policy is performed based on unsupervised learning such as trajectory VAE. However, a limitation of this approach is that the quality of the low-level policy highly depends on the quality of the dataset. To address this issue, we introduce QDOS (Quality-Diversity Offline Skill learning), a unified pipeline for robust offline-to-online learning. Our approach incorporates an Advantage-Weighted Quality-Diversity pretraining objective, which weights the skill extraction and diversity objectives by the estimated advantage of each trajectory segment. This approach allows the model to extract diverse and high-value skills. By providing robust and task-relevant skill representations, QDOS significantly improves the quality of the embedded skill space used by the low-level policy. We further integrate this with a dual dataset reuse strategy, where offline data is used both for skill pretraining and for populating the online replay buffer via pseudo-labeling. Experiments demonstrate that QDOS significantly outperforms strong baselines in structured manipulation tasks and unstructured locomotion tasks, confirming its ability to accelerate exploration and improve final returns in challenging sparse-reward domains.

📄 PDF Abstract BibTeX arXiv:2608.19684

Code (3)

Aaron617/agent-arXiv-daily ★ 10
Tavish9/awesome-daily-AI-arxiv ★ 113
arxivsub/arXivSub_daily_arxiv ★ 4

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Robust Policy Learning via Offline Skill Diffusion

2024-03-01 · Woo Kyung Kim, Minjong Yoo, Honguk Woo

Skill-based reinforcement learning (RL) approaches have shown considerable promise, especially in solving long-horizon tasks via hierarchical structures. These skills, learned task-agnostically from offline datasets, can…

DecoderImitation LearningReinforcement Learning (RL)

Offline Diversity Maximization Under Imitation Constraints

2023-07-21 · Marin Vlastelica, Jin Cheng, Georg Martius, Pavel Kolev

There has been significant recent progress in the area of unsupervised skill discovery, utilizing various information-theoretic objectives as measures of diversity. Despite these advances, challenges remain: current meth…

D4RLDiversityImitation Learning

Skill-Critic: Refining Learned Skills for Hierarchical Reinforcement Learning

2023-06-14 · Ce Hao, Catherine Weaver, Chen Tang, Kenta Kawamoto 외

Hierarchical reinforcement learning (RL) can accelerate long-horizon decision-making by temporally abstracting a policy into multiple levels. Promising results in sparse reward environments have been seen with skills, i.…

Autonomous RacingDecision MakingHierarchical Reinforcement Learningreinforcement-learning+2

Neuroevolution is a Competitive Alternative to Reinforcement Learning for Skill Discovery

2022-10-06 · Felix Chalumeau, Raphael Boige, Bryan Lim, Valentin Macé 외

Deep Reinforcement Learning (RL) has emerged as a powerful paradigm for training neural policies to solve complex control tasks. However, these policies tend to be overfit to the exact specifications of the task and envi…

Deep Reinforcement LearningDiversityreinforcement-learningReinforcement Learning (RL)

Online Damage Recovery for Physical Robots with Hierarchical Quality-Diversity

2022-10-18 · Maxime Allard, Simón C. Smith, Konstantinos Chatzilygeroudis, Bryan Lim 외

In real-world environments, robots need to be resilient to damages and robust to unforeseen scenarios. Quality-Diversity (QD) algorithms have been successfully used to make robots adapt to damages in seconds by leveragin…

Diversity