paper-with-me

홈 › Papers

Reinforcement Learning without Ground-Truth State

2019-05-20 · Xingyu Lin, Harjatin Singh Baweja, David Held

To perform robot manipulation tasks, a low-dimensional state of the environment typically needs to be estimated. However, designing a state estimator can sometimes be difficult, especially in environments with deformable objects. An alternative is to learn an end-to-end policy that maps directly from high-dimensional sensor inputs to actions. However, if this policy is trained with reinforcement learning, then without a state estimator, it is hard to specify a reward function based on high-dimensional observations. To meet this challenge, we propose a simple indicator reward function for goal-conditioned reinforcement learning: we only give a positive reward when the robot's observation exactly matches a target goal observation. We show that by relabeling the original goal with the achieved goal to obtain positive rewards (Andrychowicz et al., 2017), we can learn with the indicator reward function even in continuous state spaces. We propose two methods to further speed up convergence with indicator rewards: reward balancing and reward filtering. We show comparable performance between our method and an oracle which uses the ground-truth state for computing rewards. We show that our method can perform complex tasks in continuous state spaces such as rope manipulation from RGB-D images, without knowledge of the ground-truth state.

📄 PDF Abstract BibTeX arXiv:1905.07866

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement LearningReinforcement Learning (RL)Robot Manipulation

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

Reinforcement Learning without Ground-Truth Solutions can Improve LLMs

2026-06-25 · Yingyu Lin, Qiyue Gao, Nikki Lijing Kuang, Xunpeng Huang 외 arxiv

Reinforcement learning with verifiable rewards (RLVR) for training LLMs typically rely on ground-truth answers to assign rewards, limiting their applicability to tasks where the ground-truth solution is unknown. We intro…

Reinforcement Learning

Reinforcement Learning from Meta-Evaluation: Aligning Language Models Without Ground-Truth Labels

2026-01-29 · Micah Rentschler, Jesse Roberts arxiv

Most reinforcement learning (RL) methods for training large language models (LLMs) require ground-truth labels or task-specific verifiers, limiting scalability when correctness is ambiguous or expensive to obtain. We int…

Reinforcement Learning

PFRL: Pose-Free Reinforcement Learning for 6D Pose Estimation

2021-02-24 · CVPR 2020 6 · Jianzhun Shao, Yuhang Jiang, Gu Wang, Zhigang Li 외

6D pose estimation from a single RGB image is a challenging and vital task in computer vision. The current mainstream deep model methods resort to 2D images annotated with real-world ground-truth 6D object poses, whose c…

6D Pose EstimationPose Estimationreinforcement-learningReinforcement Learning+1

Surrogate Signals from Format and Length: Reinforcement Learning for Solving Mathematical Problems without Ground Truth Answers

2025-05-26 · Rihui Xin, Han Liu, Zecheng Wang, Yupeng Zhang 외

Large Language Models have achieved remarkable success in natural language processing tasks, with Reinforcement Learning playing a key role in adapting them to specific applications. However, obtaining ground truth answe…

Logical ReasoningMathematical Problem-Solving

Training deep learning based image denoisers from undersampled measurements without ground truth and without image prior

2018-06-04 · CVPR 2019 6 · Magauiya Zhussip, Shakarim Soltanayev, Se Young Chun

Compressive sensing is a method to recover the original image from undersampled measurements. In order to overcome the ill-posedness of this inverse problem, image priors are used such as sparsity in the wavelet domain, …

Compressive SensingDeep Learning