Localized Dynamics-Aware Domain Adaption for Off-Dynamics Offline Reinforcement Learning
Off-dynamics offline reinforcement learning (RL) aims to learn a policy for a target domain using limited target data and abundant source data collected under different transition dynamics. Existing methods typically address dynamics mismatch either globally over the state space or via pointwise data filtering; these approaches can miss localized cross-domain similarities or incur high computational cost. We propose Localized Dynamics-Aware Domain Adaptation (LoDADA), which exploits localized dynamics mismatch to better reuse source data. LoDADA clusters transitions from source and target datasets and estimates cluster-level dynamics discrepancy via domain discrimination. Source transitions from clusters with small discrepancy are retained, while those from clusters with large discrepancy are filtered out. This yields a fine-grained and scalable data selection strategy that avoids overly coarse global assumptions and expensive per-sample filtering. We provide theoretical insights and extensive experiments across environments with diverse global and local dynamics shifts. Results show that LoDADA consistently outperforms state-of-the-art off-dynamics offline RL methods by better leveraging localized distribution mismatch.
Code (0)
등록된 구현이 없습니다.
Tasks
Reinforcement LearningDomain AdaptationOffline RLSimilar Papers 제목 키워드 기반
Zero-Shot Compositional Policy Learning via Language Grounding
Despite recent breakthroughs in reinforcement learning (RL) and imitation learning (IL), existing algorithms fail to generalize beyond the training environments. In reality, humans can adapt to new tasks quickly by lever…
DescriptiveDomain AdaptationGrounded language learningImitation Learning+4OMPO: A Unified Framework for RL under Policy and Dynamics Shifts
Training reinforcement learning policies using environment interaction data collected from varying policies or dynamics presents a fundamental challenge. Existing works often overlook the distribution discrepancies induc…
Domain AdaptationOpenAI GymG-PARC: Graph-Physics Aware Recurrent Convolutional Neural Networks for Spatiotemporal Dynamics on Unstructured Meshes
Physics-aware recurrent convolutional networks (PARC) have demonstrated strong performance in predicting nonlinear spatiotemporal dynamics by embedding differential operators directly into the computational graph of a ne…
Integrating Egocentric Localization for More Realistic Point-Goal Navigation Agents
Recent work has presented embodied agents that can navigate to point-goal targets in novel indoor environments with near-perfect accuracy. However, these agents are equipped with idealized sensors for localization and ta…
NavigateRobot NavigationVisual OdometryLocalized modulated wave solutions in diffusive glucose-insulin systems
We investigate intercellular insulin dynamics in an array of diffusively coupled pancreatic islet \b{eta}-cells. The cells are connected via gap junction coupling, where nearest neighbor interactions are included. Throug…