paper-with-me

Papers D4RL

“D4RL” 태그가 달린 논문 226편 · 필터 해제

From Novelty to Imitation: Self-Distilled Rewards for Offline Reinforcement Learning

2025-07-17 · Gaurav Chaudhary, Laxmidhar Behera

Offline Reinforcement Learning (RL) aims to learn effective policies from a static dataset without requiring further agent-environment interactions. However, its practical adoption is often hindered by the need for expli…

D4RLOffline RLreinforcement-learningReinforcement Learning+1

Accelerating Residual Reinforcement Learning with Uncertainty Estimation

2025-06-21 · Lakshita Dodeja, Karl Schmeckpeper, Shivam Vats, Thomas Weng 외

Residual Reinforcement Learning (RL) is a popular approach for adapting pretrained policies by learning a lightweight residual policy that provides corrective actions. While Residual RL is more sample-efficient than fine…

D4RLreinforcement-learningReinforcement LearningReinforcement Learning (RL)

CAWR: Corruption-Averse Advantage-Weighted Regression for Robust Policy Optimization

2025-06-18 · Ranting Hu

Offline reinforcement learning (offline RL) algorithms often require additional constraints or penalty terms to address distribution shift issues, such as adding implicit or explicit policy constraints during policy opti…

D4RLOffline RLregression

MOORL: A Framework for Integrating Offline-Online Reinforcement Learning

2025-06-11 · Gaurav Chaudhary, Wassim Uddin Mondal, Laxmidhar Behera

Sample efficiency and exploration remain critical challenges in Deep Reinforcement Learning (DRL), particularly in complex domains. Offline RL, which enables agents to learn optimal policies from static, pre-collected da…

D4RLDeep Reinforcement LearningEfficient ExplorationOffline RL+2

Offline RL with Smooth OOD Generalization in Convex Hull and its Neighborhood

2025-06-10 · Qingmao Yao, Zhichao Lei, Tianyuan Chen, Ziyue Yuan 외

Offline Reinforcement Learning (RL) struggles with distributional shifts, leading to the $Q$-value overestimation for out-of-distribution (OOD) actions. Existing methods address this issue by imposing constraints; howeve…

Computational EfficiencyD4RLOffline RLReinforcement Learning (RL)

Policy-Based Trajectory Clustering in Offline Reinforcement Learning

2025-06-10 · Hao Hu, Xinqi Wang, Simon Shaolei Du

We introduce a novel task of clustering trajectories from offline reinforcement learning (RL) datasets, where each cluster center represents the policy that generated its trajectories. By leveraging the connection betwee…

ClusteringD4RLOffline RLreinforcement-learning+4

STITCH-OPE: Trajectory Stitching with Guided Diffusion for Off-Policy Evaluation

2025-05-27 · Hossein Goli, Michael Gimelfarb, Nathan Samuel de Lara, Haruki Nishimura 외

Off-policy evaluation (OPE) estimates the performance of a target policy using offline data collected from a behavior policy, and is crucial in domains such as robotics or healthcare where direct interaction with the env…

D4RLDenoisingOff-policy evaluationOpenAI Gym

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL

2025-05-26 · Qin-Wen Luo, Ming-Kun Xie, Ye-Wen Wang, Sheng-Jun Huang

Offline reinforcement learning (RL) aims to learn an effective policy from a static dataset. To alleviate extrapolation errors, existing studies often uniformly regularize the value function or policy updates across all …

D4RLOffline RLReinforcement Learning (RL)

Temporal Distance-aware Transition Augmentation for Offline Model-based Reinforcement Learning

2025-05-19 · Dongsu Lee, Minhae Kwon

The goal of offline reinforcement learning (RL) is to extract a high-performance policy from the fixed datasets, minimizing performance degradation due to out-of-distribution (OOD) samples. Offline model-based RL (MBRL) …

D4RLModel-based Reinforcement LearningReinforcement Learning (RL)

Policy-Driven World Model Adaptation for Robust Offline Model-based Reinforcement Learning

2025-05-19 · Jiayu Chen, Aravind Venugopal, Jeff Schneider

Offline reinforcement learning (RL) offers a powerful paradigm for data-driven control. Compared to model-free approaches, offline model-based RL (MBRL) explicitly learns a world model from a static dataset and uses it a…

D4RLmodelModel-based Reinforcement LearningMuJoCo+1

Imagination-Limited Q-Learning for Offline Reinforcement Learning

2025-05-18 · Wenhui Liu, Zhijian Wu, JingChao Wang, Dingjiang Huang 외

Offline reinforcement learning seeks to derive improved policies entirely from historical data but often struggles with over-optimistic value estimates for out-of-distribution (OOD) actions. This issue is typically mitig…

D4RLQ-Learningreinforcement-learningReinforcement Learning

Beyond the Known: Decision Making with Counterfactual Reasoning Decision Transformer

2025-05-14 · Minh Hoang Nguyen, Linh Le Pham Van, Thommen George Karimpanal, Sunil Gupta 외

Decision Transformers (DT) play a crucial role in modern reinforcement learning, leveraging offline datasets to achieve impressive results across various domains. However, DT requires high-quality, comprehensive data to …

counterfactualCounterfactual ReasoningD4RLDecision Making+2

Pretraining a Shared Q-Network for Data-Efficient Offline Reinforcement Learning

2025-05-09 · Jongchan Park, MinGyu Park, Donghwan Lee

Offline reinforcement learning (RL) aims to learn a policy from a static dataset without further interactions with the environment. Collecting sufficiently large datasets for offline RL is exhausting since this data coll…

D4RLOffline RLReinforcement Learning (RL)

Taming OOD Actions for Offline Reinforcement Learning: An Advantage-Based Approach

2025-05-08 · Xuyang Chen, Keyu Yan, Lin Zhao

Offline reinforcement learning (RL) aims to learn decision-making policies from fixed datasets without online interactions, providing a practical solution where online data collection is expensive or risky. However, offl…

D4RLDecision MakingOffline RLReinforcement Learning (RL)

Analytic Energy-Guided Policy Optimization for Offline Reinforcement Learning

2025-05-03 · Jifeng Hu, Sili Huang, Zhejian Yang, Shengchao Hu 외

Conditional decision generation with diffusion models has shown powerful competitiveness in reinforcement learning (RL). Recent studies reveal the relation between energy-function-guidance diffusion models and constraine…

D4RLOffline RLreinforcement-learningReinforcement Learning+1

Directly Forecasting Belief for Reinforcement Learning with Delays

2025-05-01 · Qingyuan Wu, Yuhui Wang, Simon Sinong Zhan, YiXuan Wang 외

Reinforcement learning (RL) with delays is challenging as sensory perceptions lag behind the actual events: the RL agent needs to estimate the real state of its environment based on past observations. State-of-the-art (S…

D4RLMuJoCoreinforcement-learningReinforcement Learning+1

An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning

2025-04-17 · Haoran Xu, Shuozhe Li, Harshit Sikchi, Scott Niekum 외

We introduce Iterative Dual Reinforcement Learning (IDRL), a new method that takes an optimal discriminator-weighted imitation view of solving RL. Our method is motivated by a simple experiment in which we find training …

D4RLreinforcement-learningReinforcement Learning

VIPO: Value Function Inconsistency Penalized Offline Reinforcement Learning

2025-04-16 · Xuyang Chen, GuoJian Wang, Keyu Yan, Lin Zhao

Offline reinforcement learning (RL) learns effective policies from pre-collected datasets, offering a practical solution for applications where online interactions are risky or costly. Model-based approaches are particul…

D4RLOffline RLreinforcement-learningReinforcement Learning+1

Decision SpikeFormer: Spike-Driven Transformer for Decision Making

2025-04-04 · CVPR 2025 1 · Wei Huang, Qinying Gu, Nanyang Ye

Offline reinforcement learning (RL) enables policy training solely on pre-collected data, avoiding direct environment interaction - a crucial benefit for energy-constrained embodied AI applications. Although Artificial N…

D4RLDecision MakingOffline RLReinforcement Learning (RL)

Model-Based Offline Reinforcement Learning with Adversarial Data Augmentation

2025-03-26 · Hongye Cao, Fan Feng, Jing Huo, Shangdong Yang 외

Model-based offline Reinforcement Learning (RL) constructs environment models from offline datasets to perform conservative policy optimization. Existing approaches focus on learning state transitions through ensemble mo…

D4RLData AugmentationOffline RLreinforcement-learning+2
1–20 / 226 다음 →