paper-with-me

홈 › Papers

SLOPE: Optimistic Potential Landscape Shaping for Model-based Reinforcement Learning

2026-02-03 · Yao-Hui Li, Zeyu Wang, Xin Li, Wei Pang, Yingfang Yuan, Zhengkun Chen, Boya Zhang, Riashat Islam, Alex Lamb, Yonggang Zhang arxiv

Model-based reinforcement learning (MBRL) is sample-efficient but struggles in sparse reward settings. A critical bottleneck arises from the lack of informative gradients in sparse settings, where standard reward models often yield flat landscapes that struggle to guide planning. To address this challenge, we propose Shaping Landscapes with Optimistic Potential Estimates (SLOPE), a novel framework that shifts reward modeling from predicting sparse scalars to constructing informative potential landscapes. SLOPE employs optimistic distributional regression to estimate high-confidence upper bounds, which amplifies rare success signals and ensures sufficient exploration gradients. Evaluations on 30+ tasks across 5 benchmarks and real-world robotic deployments, demonstrate that SLOPE consistently outperforms leading baselines in fully sparse, semi-sparse, and dense rewards.

📄 PDF Abstract BibTeX arXiv:2602.03201

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Moving mountains: grazing agents drive terracette formation on steep hillslopes

2025-04-23 · Benjamin Seleb, Atanu Chatterjee, Saad Bhamla

Terracettes, striking, step-like landforms that stripe steep, vegetated hillslopes, have puzzled scientists for more than a century. Competing hypotheses invoke either slow mass-wasting or the relentless trampling of gra…

On Surprising Effects of Risk-Aware Domain Randomization for Contact-Rich Sampling-based Predictive Control

2026-05-05 · Sergio A. Esteban, Junheng Li, Vince Kurtz, Aaron D. Ames arxiv

Domain randomization (DR) is widely used in policy learning to improve robustness to modeling error, but remains underexplored in contact-rich sampling-based predictive control (SPC), where rollout quality is highly sens…

Optimistic Curiosity Exploration and Conservative Exploitation with Linear Reward Shaping

2022-09-15 · Hao Sun, Lei Han, Rui Yang, Xiaoteng Ma 외

In this work, we study the simple yet universally applicable case of reward shaping in value-based Deep Reinforcement Learning (DRL). We show that reward shifting in the form of the linear transformation is equivalent to…

continuous-controlContinuous ControlDeep Reinforcement LearningOffline RL

Automatic delineation of geomorphological slope units with r.slopeunits v1.0 and their optimization for landslide susceptibility modeling

2016-11-09 · 21/06 2016 11 · Massimiliano Alvioli, Ivan Marchesini, Paola Reichenbach, Mauro Rossi 외

Automatic subdivision of landscapes into terrain units remains a challenge. Slope units are terrain units bounded by drainage and divide lines, but their use in hydrological and geomorphological studies is limited bec…

Magnetic Field-Based Reward Shaping for Goal-Conditioned Reinforcement Learning

2023-07-16 · Hongyu Ding, Yuanze Tang, Qing Wu, Bo wang 외

Goal-conditioned reinforcement learning (RL) is an interesting extension of the traditional RL framework, where the dynamic environment and reward sparsity can cause conventional learning algorithms to fail. Reward shapi…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)