paper-with-me

홈 › Papers

Learning Contraction Policies from Offline Data

2021-12-11 · Navid Rezazadeh, Maxwell Kolarich, Solmaz S. Kia, Negar Mehr

This paper proposes a data-driven method for learning convergent control policies from offline data using Contraction theory. Contraction theory enables constructing a policy that makes the closed-loop system trajectories inherently convergent towards a unique trajectory. At the technical level, identifying the contraction metric, which is the distance metric with respect to which a robot's trajectories exhibit contraction is often non-trivial. We propose to jointly learn the control policy and its corresponding contraction metric while enforcing contraction. To achieve this, we learn an implicit dynamics model of the robotic system from an offline data set consisting of the robot's state and input trajectories. Using this learned dynamics model, we propose a data augmentation algorithm for learning contraction policies. We randomly generate samples in the state-space and propagate them forward in time through the learned dynamics model to generate auxiliary sample trajectories. We then learn both the control policy and the contraction metric such that the distance between the trajectories from the offline data set and our generated auxiliary sample trajectories decreases over time. We evaluate the performance of our proposed framework on simulated robotic goal-reaching tasks and demonstrate that enforcing contraction results in faster convergence and greater robustness of the learned policy.

📄 PDF Abstract BibTeX arXiv:2112.05911

Code (0)

등록된 구현이 없습니다.

Tasks

Data Augmentation

Similar Papers 제목 키워드 기반

Contractive Diffusion Policies: Robust Action Diffusion via Contractive Score-Based Sampling with Differential Equations

2026-01-02 · Amin Abyaneh, Charlotte Morissette, Mohamad H. Danesh, Anas El Houssaini 외 arxiv

Diffusion policies have emerged as powerful generative models for offline policy learning, whose sampling process can be rigorously characterized by a score function guiding a stochastic differential equation (SDE). Howe…

Continuous ControlImage Generation

Residuals-based Offline Reinforcement Learning

2026-04-01 · Qing Zhu, Xian Yu arxiv

Offline reinforcement learning (RL) has received increasing attention for learning policies from previously collected data without interaction with the real environment, which is particularly important in high-stakes app…

Reinforcement LearningOffline RL

Offline Policy Evaluation for Manipulation Policies via Discounted Liveness Formulation

2026-05-12 · Hao Wang, Joshua Bowden, Colton Crosby, Somil Bansal arxiv

Policy evaluation is a fundamental component of the development and deployment pipeline for robotic policies. In modern manipulation systems, this problem is particularly challenging: rewards are often sparse, task progr…

ContractionPPO: Certified Reinforcement Learning via Differentiable Contraction Layers

2026-03-20 · Vrushabh Zinage, Narek Harutyunyan, Eric Verheyden, Fred Y. Hadaegh 외 arxiv

Legged locomotion in unstructured environments demands not only high-performance control policies but also formal guarantees to ensure robustness under perturbations. Control methods often require carefully designed refe…

Reinforcement Learning

Guardian: Decoupling Exploration from Safety in Reinforcement Learning

2025-10-26 · Kaitong Cai, Jusheng Zhang, Jing Yang, Keze Wang arxiv

Hybrid offline--online reinforcement learning (O2O RL) promises both sample efficiency and robust exploration, but suffers from instability due to distribution shift between offline and online data. We introduce RLPD-GX,…

Reinforcement Learning