paper-with-me

D4RL

1개 벤치마크 · 논문 226편 · 이 태스크의 논문 보기 →

Benchmarks

D4RL

결과 9개

Most implemented

Reformer: The Efficient Transformer

2020-01-13 · 구현 10개

Rethinking Attention with Performers

2020-09-30 · 구현 7개

Papers

From Novelty to Imitation: Self-Distilled Rewards for Offline Reinforcement Learning

2025-07-17 · Gaurav Chaudhary, Laxmidhar Behera

Offline Reinforcement Learning (RL) aims to learn effective policies from a static dataset without requiring further agent-environment interactions. However, its practical adoption is often hindered by the need for expli…

D4RLOffline RLreinforcement-learningReinforcement Learning+1

Accelerating Residual Reinforcement Learning with Uncertainty Estimation

2025-06-21 · Lakshita Dodeja, Karl Schmeckpeper, Shivam Vats, Thomas Weng 외

Residual Reinforcement Learning (RL) is a popular approach for adapting pretrained policies by learning a lightweight residual policy that provides corrective actions. While Residual RL is more sample-efficient than fine…

D4RLreinforcement-learningReinforcement LearningReinforcement Learning (RL)

CAWR: Corruption-Averse Advantage-Weighted Regression for Robust Policy Optimization

2025-06-18 · Ranting Hu

Offline reinforcement learning (offline RL) algorithms often require additional constraints or penalty terms to address distribution shift issues, such as adding implicit or explicit policy constraints during policy opti…

D4RLOffline RLregression

MOORL: A Framework for Integrating Offline-Online Reinforcement Learning

2025-06-11 · Gaurav Chaudhary, Wassim Uddin Mondal, Laxmidhar Behera

Sample efficiency and exploration remain critical challenges in Deep Reinforcement Learning (DRL), particularly in complex domains. Offline RL, which enables agents to learn optimal policies from static, pre-collected da…

D4RLDeep Reinforcement LearningEfficient ExplorationOffline RL+2

Offline RL with Smooth OOD Generalization in Convex Hull and its Neighborhood

2025-06-10 · Qingmao Yao, Zhichao Lei, Tianyuan Chen, Ziyue Yuan 외

Offline Reinforcement Learning (RL) struggles with distributional shifts, leading to the $Q$-value overestimation for out-of-distribution (OOD) actions. Existing methods address this issue by imposing constraints; howeve…

Computational EfficiencyD4RLOffline RLReinforcement Learning (RL)

Policy-Based Trajectory Clustering in Offline Reinforcement Learning

2025-06-10 · Hao Hu, Xinqi Wang, Simon Shaolei Du

We introduce a novel task of clustering trajectories from offline reinforcement learning (RL) datasets, where each cluster center represents the policy that generated its trajectories. By leveraging the connection betwee…

ClusteringD4RLOffline RLreinforcement-learning+4

전체 226편 보기 →