paper-with-me

홈 › Papers

Policy Learning Using Weak Supervision

2020-10-05 · NeurIPS 2021 12 · Jingkang Wang, Hongyi Guo, Zhaowei Zhu, Yang Liu

Most existing policy learning solutions require the learning agents to receive high-quality supervision signals such as well-designed rewards in reinforcement learning (RL) or high-quality expert demonstrations in behavioral cloning (BC). These quality supervisions are usually infeasible or prohibitively expensive to obtain in practice. We aim for a unified framework that leverages the available cheap weak supervisions to perform policy learning efficiently. To handle this problem, we treat the "weak supervision" as imperfect information coming from a peer agent, and evaluate the learning agent's policy based on a "correlated agreement" with the peer agent's policy (instead of simple agreements). Our approach explicitly punishes a policy for overfitting to the weak supervision. In addition to theoretical guarantees, extensive evaluations on tasks including RL with noisy rewards, BC with weak demonstrations, and standard policy co-training show that our method leads to substantial performance improvements, especially when the complexity or the noise of the learning environments is high.

📄 PDF Abstract BibTeX arXiv:2010.01748

Code (1)

wangjksjtu/peerpl 공식 구현

Tasks

Reinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Weak Supervision in Analysis of News: Application to Economic Policy Uncertainty

2022-08-10 · Paul Trust, Ahmed Zahran, Rosane Minghim

The need for timely data analysis for economic decisions has prompted most economists and policy makers to search for non-traditional supplementary sources of data. In that context, text data is being explored to enrich …

Articles

Weak Critics Make Strong Learners: On-Policy Critique Distillation for Scalable Oversight

2026-05-29 · Can Jin, Jiakang Li, Rui Wu, Eddy Zhang 외 arxiv

As large language models become stronger, weak supervisors may fail to provide reliable labels, preferences, or final judgments for complex outputs, limiting both weak-to-strong generalization and scalable oversight. We …

Weakly Supervised Reinforcement Learning for Autonomous Highway Driving via Virtual Safety Cages

2021-03-17 · Sampo Kuutti, Richard Bowden, Saber Fallah

The use of neural networks and reinforcement learning has become increasingly popular in autonomous vehicle control. However, the opaqueness of the resulting control policies presents a significant barrier to deploying n…

Autonomous Vehiclesreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Extreme Region Policy Distillation

2026-05-25 · Changyu Chen, Xiting Wang, Rui Yan arxiv

Reinforcement learning for large language models faces a fundamental trade-off between sample efficiency and asymptotic performance: strictly on-policy methods discard trajectories after a single update, while off-policy…

Reinforcement LearningMathematical Reasoning

Weak-to-Strong Generalization via Direct On-Policy Distillation

2026-07-06 · Shiyuan Feng, Huan-ang Gao, Haohan Chi, Hanlin Wu 외 arxiv

Reinforcement learning with verifiable rewards (RLVR) is a powerful recipe for improving language-model reasoning, but it is expensive to repeat on every new strong model because the target model must generate many rollo…

Reinforcement Learning