paper-with-me

홈 › Papers

Renaissance Robot: Optimal Transport Policy Fusion for Learning Diverse Skills

2022-07-03 · Julia Tan, Ransalu Senanayake, Fabio Ramos

Deep reinforcement learning (RL) is a promising approach to solving complex robotics problems. However, the process of learning through trial-and-error interactions is often highly time-consuming, despite recent advancements in RL algorithms. Additionally, the success of RL is critically dependent on how well the reward-shaping function suits the task, which is also time-consuming to design. As agents trained on a variety of robotics problems continue to proliferate, the ability to reuse their valuable learning for new domains becomes increasingly significant. In this paper, we propose a post-hoc technique for policy fusion using Optimal Transport theory as a robust means of consolidating the knowledge of multiple agents that have been trained on distinct scenarios. We further demonstrate that this provides an improved weights initialisation of the neural network policy for learning new tasks, requiring less time and computational resources than either retraining the parent policies or training a new policy from scratch. Ultimately, our results on diverse agents commonly used in deep RL show that specialised knowledge can be unified into a "Renaissance agent", allowing for quicker learning of new skills.

📄 PDF Abstract BibTeX arXiv:2207.00978

Code (1)

julia-lina-tan/ot-policy-fusion 공식 구현

Tasks

Deep Reinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Optimal Transport Q-Learning for Flow Policy Steering and Acceleration

2026-07-07 · Andreas Sochopoulos, Esmeralda S. Whitammer, Nikolaos Tsagkas, João Moura 외 arxiv

Diffusion and flow policies have recently demonstrated remarkable performance in robotic applications by accurately capturing multimodal robot trajectory distributions, especially in the context of vision language action…

Score-Based Diffusion Policy Compatible with Reinforcement Learning via Optimal Transport

2025-02-18 · Mingyang Sun, Pengxiang Ding, Weinan Zhang, Donglin Wang

Diffusion policies have shown promise in learning complex behaviors from demonstrations, particularly for tasks requiring precise control and long-term planning. However, they face challenges in robustness when encounter…

Imitation Learning

Ambient Diffusion Policy: Imitation Learning from Suboptimal Data in Robotics

2026-06-10 · Adam Wei, Nicholas Pfaff, Thomas Cohn, Arif Kerem Dayı 외 arxiv

We propose Ambient Diffusion Policy, a simple and principled method for imitation learning from suboptimal data in robotics. High-quality, task-specific robot data is expensive and time-consuming to collect, while subopt…

Hierarchical Policy Blending As Optimal Transport

2022-12-04 · An T. Le, Kay Hansel, Jan Peters, Georgia Chalvatzaki

We present hierarchical policy blending as optimal transport (HiPBOT). HiPBOT hierarchically adjusts the weights of low-level reactive expert policies of different agents by adding a look-ahead planning layer on the para…

Robot Policy Learning with Temporal Optimal Transport Reward

2024-10-29 · Yuwei Fu, Haichao Zhang, Di wu, Wei Xu 외

Reward specification is one of the most tricky problems in Reinforcement Learning, which usually requires tedious hand engineering in practice. One promising approach to tackle this challenge is to adopt existing expert …