paper-with-me

홈 › Papers

Conservative Offline Robot Policy Learning via Posterior-Transition Reweighting

2026-03-17 · Wanpeng Zhang, Hao Luo, Sipeng Zheng, Yicheng Feng, Haiweng Xu, Ziheng Xi, Chaoyi Xu, Haoqi Yuan, Zongqing Lu arxiv

Offline post-training adapts a pretrained robot policy to a target dataset by supervised regression on recorded actions. In practice, robot datasets are heterogeneous: they mix embodiments, camera setups, and demonstrations of varying quality, so many trajectories reflect recovery behavior, inconsistent operator skill, or weakly informative supervision. Uniform post-training gives equal credit to all samples and can therefore average over conflicting or low-attribution data. We propose Posterior-Transition Reweighting (PTR), a reward-free and conservative post-training method that decides how much each training sample should influence the supervised update. For each sample, PTR encodes the observed post-action consequence as a latent target, inserts it into a candidate pool of mismatched targets, and uses a separate transition scorer to estimate a softmax identification posterior over target indices. The posterior-to-uniform ratio defines the PTR score, which is converted into a clipped-and-mixed weight and applied to the original action objective through self-normalized weighted regression. This construction requires no tractable policy likelihood and is compatible with both diffusion and flow-matching action heads. Rather than uniformly trusting all recorded supervision, PTR reallocates credit according to how attributable each sample's post-action consequence is under the current representation, improving conservative offline adaptation to heterogeneous robot data.

📄 PDF Abstract BibTeX arXiv:2603.16542

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Bayesian Conservative Policy Optimization (BCPO): A Novel Uncertainty-Calibrated Offline Reinforcement Learning with Credible Lower Bounds

2026-03-06 · Debashis Chatterjee arxiv

Offline reinforcement learning (RL) aims to learn decision policies from a fixed batch of logged transitions, without additional environment interaction. Despite remarkable empirical progress, offline RL remains fragile …

Reinforcement LearningOffline RL

Contextual Conservative Q-Learning for Offline Reinforcement Learning

2023-01-03 · Ke Jiang, Jiayu Yao, Xiaoyang Tan

Offline reinforcement learning learns an effective policy on offline datasets without online interaction, and it attracts persistent research attention due to its potential of practical application. However, extrapolatio…

MuJoCoQ-Learningreinforcement-learningReinforcement Learning+1

ACL-QL: Adaptive Conservative Level in Q-Learning for Offline Reinforcement Learning

2024-12-22 · Kun Wu, Yinuo Zhao, Zhiyuan Xu, Zhengping Che 외

Offline Reinforcement Learning (RL), which operates solely on static datasets without further interactions with the environment, provides an appealing alternative to learning a safe and promising control policy. The prev…

D4RLQ-LearningReinforcement Learning (RL)

Environment Transformer and Policy Optimization for Model-Based Offline Reinforcement Learning

2023-03-07 · Pengqin Wang, Meixin Zhu, Shaojie Shen

Interacting with the actual environment to acquire data is often costly and time-consuming in robotic tasks. Model-based offline reinforcement learning (RL) provides a feasible solution. On the one hand, it eliminates th…

Continuous ControlOffline RLQ-Learningreinforcement-learning+2

Robust Reinforcement Learning for Continuous Control with Model Misspecification

2019-06-18 · ICLR 2020 1 · Daniel J. Mankowitz, Nir Levine, Rae Jeong, Yuanyuan Shi 외

We provide a framework for incorporating robustness -- to perturbations in the transition dynamics which we refer to as model misspecification -- into continuous control Reinforcement Learning (RL) algorithms. We specifi…

continuous-controlContinuous ControlMuJoCoreinforcement-learning+2