paper-with-me

Papers

Diffusion-Guided Uncertainty-Aware Delayed Policy Optimization

2026-07-06 · Junqi Tu, Zejiao Liu, Fangfei Li, Yang Tang arxiv

Reinforcement learning in real world environments often suffers from severe performance degradation due to delayed feedback. Existing approaches typically mitigate performance degradation caused by observation delays by constructing augmented states or predicting the true states. However, these methods often overlook the inherent discrepancy between delayed state and true states induced by stochastic MDP. We theoretically prove the existence of such a discrepancy and show that it leads to the degradation of the optimal policy. To address this challenge, we propose Diffusion Guided Uncertainty Aware Delayed Policy Optimization (DUPO). Our method explicitly models the relationship between delayed state message and the current state using a diffusion model, and leverages the resulting discrepancy estimates to weight delayed policies. Extensive experiments on continuous robotic control tasks with multiple stochastic delays demonstrate that DUPO consistently outperforms existing methods and remains effective even under long and random delay scenarios.

📄 PDF Abstract BibTeX arXiv:2607.05064

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Uncertainty-Aware LLM-Guided Policy Shaping for Sparse-Reward Reinforcement Learning

2026-06-04 · Ujjwal Bhatta, Utsabi Dangol, Sumaly Bajracharya, Rodrigue Rizk 외 arxiv

Sparse rewards and heterogeneous task sequences remain persistent challenges in Reinforcement Learning (RL), often resulting in slow convergence, weak generalization, and inefficient exploration. We propose Uncertainty-A…

Reinforcement Learning

QUSR: Quality-Aware and Uncertainty-Guided Image Super-Resolution Diffusion Model

2026-03-10 · Junjie Yin, Jiaju Li, Hanfa Xing arxiv

Diffusion-based image super-resolution (ISR) has shown strong potential, but it still struggles in real-world scenarios where degradations are unknown and spatially non-uniform, often resulting in lost details or visual …

Image Super-Resolution

Regret-Aware Policy Optimization: Environment-Level Memory for Replay Suppression under Delayed Harm

2026-04-08 · Prakul Sunil Hiremath arxiv

Safety in reinforcement learning (RL) is typically enforced through objective shaping while keeping environment dynamics stationary with respect to observable state-action pairs. Under delayed harm, this can lead to repl…

Reinforcement Learning

Guided Uncertainty-Aware Policy Optimization: Combining Learning and Model-Based Strategies for Sample-Efficient Policy Learning

2020-05-21 · Michelle A. Lee, Carlos Florensa, Jonathan Tremblay, Nathan Ratliff 외

Traditional robotic approaches rely on an accurate model of the environment, a detailed description of how to perform the task, and a robust perception system to keep track of the current state. On the other hand, reinfo…

Uncertainty-Aware Multi-Objective Reinforcement Learning-Guided Diffusion Models for 3D De Novo Molecular Design

2025-10-24 · Lianghong Chen, Dongkyu Eugene Kim, Mike Domaratzki, Pingzhao Hu arxiv

Designing de novo 3D molecules with desirable properties remains a fundamental challenge in drug discovery and molecular engineering. While diffusion models have demonstrated remarkable capabilities in generating high-qu…

Reinforcement LearningDrug Discovery