paper-with-me

홈 › Papers

Bias-Aware Heapified Policy for Active Learning

2019-11-18 · Wen-Yen Chang, Wen-Huan Chiang, Shao-Hao Lu, Tingfan Wu, Min Sun

The data efficiency of learning-based algorithms is more and more important since high-quality and clean data is expensive as well as hard to collect. In order to achieve high model performance with the least number of samples, active learning is a technique that queries the most important subset of data from the original dataset. In active learning domain, one of the mainstream research is the heuristic uncertainty-based method which is useful for the learning-based system. Recently, a few works propose to apply policy reinforcement learning (PRL) for querying important data. It seems more general than heuristic uncertainty-based method owing that PRL method depends on data feature which is reliable than human prior. However, there have two problems - sample inefficiency of policy learning and overconfidence, when applying PRL on active learning. To be more precise, sample inefficiency of policy learning occurs when sampling within a large action space, in the meanwhile, class imbalance can lead to the overconfidence. In this paper, we propose a bias-aware policy network called Heapified Active Learning (HAL), which prevents overconfidence, and improves sample efficiency of policy learning by heapified structure without ignoring global inforamtion(overview of the whole unlabeled set). In our experiment, HAL outperforms other baseline methods on MNIST dataset and duplicated MNIST. Last but not least, we investigate the generalization of the HAL policy learned on MNIST dataset by directly applying it on MNIST-M. We show that the agent can generalize and outperform directly-learned policy under constrained labeled sets.

📄 PDF Abstract BibTeX arXiv:1911.07574

Code (0)

등록된 구현이 없습니다.

Tasks

Active LearningReinforcement Learning

Similar Papers 제목 키워드 기반

Policy-Aware Unbiased Learning to Rank for Top-k Rankings

2020-05-18 · Harrie Oosterhuis, Maarten de Rijke

Counterfactual Learning to Rank (LTR) methods optimize ranking systems using logged user interactions that contain interaction biases. Existing methods are only unbiased if users are presented with all relevant items in …

counterfactualLearning-To-RankRetrieval

Retroactive Advantage Correction: Closed-Form V-Trace Bias Correction for Delay-Aware RLHF

2026-06-25 · Arnav Raj arxiv

Reinforcement learning from human feedback (RLHF) in production does not always have a synchronous reward signal. Code-execution verifiers, slow judge ensembles, and queued human review can return several gradient steps …

Reinforcement Learning

The Hidden Bias of Process Reward Models:PRISM for Rewarding the Right Reasoning

2026-06-08 · Aakriti Agrawal, Souradip Chakraborty, Armin Saghafian, Nihal Sharma 외 arxiv

Process Reward Models (PRMs) improve credit assignment for reasoning by providing step-level feedback. However, we identify a hidden bias in PRMs caused by severe imbalance in step-level training data. Standard cross-ent…

Look Before Leap: Look-Ahead Planning with Uncertainty in Reinforcement Learning

2025-03-26 · Yongshuai Liu, Xin Liu

Model-based reinforcement learning (MBRL) has demonstrated superior sample efficiency compared to model-free reinforcement learning (MFRL). However, the presence of inaccurate models can introduce biases during policy le…

Atari GamesModel-based Reinforcement Learning

Robust Estimation and Inference in Panels with Interactive Fixed Effects

2022-10-13 · Timothy B. Armstrong, Martin Weidner, Andrei Zeleneev

We consider estimation and inference for a regression coefficient in panels with interactive fixed effects (i.e., with a factor structure). We demonstrate that existing estimators and confidence intervals (CIs) can be he…

valid