paper-with-me

Papers

Accidental exploration through value predictors

· Tomasz Kisielewski, Damian Leśniak, Maia Pasek

Infinite length of trajectories is an almost universal assumption in the theoretical foundations of reinforcement learning. In practice learning occurs on finite trajectories. In this paper we examine a specific result of this disparity, namely a strong bias of the time-bounded Every-visit Monte Carlo value estimator. This manifests as a vastly different learning dynamic for algorithms that use value predictors, including encouraging or discouraging exploration. We investigate these claims theoretically for a one dimensional random walk, and empirically on a number of simple environments. We use GAE as an algorithm involving a value predictor and evolution strategies as a reference point.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

You Must Have Clicked on this Ad by Mistake! Data-Driven Identification of Accidental Clicks on Mobile Ads with Applications to Advertiser Cost Discounting and Click-Through Rate Prediction

2018-04-03 · Gabriele Tolomei, Mounia Lalmas, Ayman Farahat, Andrew Haines

In the cost per click (CPC) pricing model, an advertiser pays an ad network only when a user clicks on an ad; in turn, the ad network gives a share of that revenue to the publisher where the ad was impressed. Still, adve…

Click-Through Rate Prediction

Learning to Generalize from Sparse and Underspecified Rewards

2019-02-19 · Rishabh Agarwal, Chen Liang, Dale Schuurmans, Mohammad Norouzi

We consider the problem of learning from sparse and underspecified rewards, where an agent receives a complex input, such as a natural language instruction, and needs to generate a complex response, such as an action seq…

Bayesian OptimizationSemantic Parsing

Class-Aware Reinforcement Learning for Counterfactual Explanation Generation

2026-07-30 · Muhammad Adil Saleem, Syed Ali Raza, Mary-Anne Williams arxiv

Counterfactual explanations (CFEs) enhance the interpretability of black-box models by generating alternative instances with adjusted feature values that achieve a contrastive outcome. Reinforcement learning (RL) offers …

Explanation GenerationReinforcement Learning

Unbiased Filtering Of Accidental Clicks in Verizon Media Native Advertising

2023-12-08 · Yohay Kaplan, Naama Krasne, Alex Shtoff, Oren Somekh

Verizon Media (VZM) native advertising is one of VZM largest and fastest growing businesses, reaching a run-rate of several hundred million USDs in the past year. Driving the VZM native models that are used to predict ev…

Collaborative Filtering

Training conformal predictors

2020-05-14 · Nicolo Colombo, Vladimir Vovk

Efficiency criteria for conformal prediction, such as \emph{observed fuzziness} (i.e., the sum of p-values associated with false labels), are commonly used to \emph{evaluate} the performance of given conformal predictors…

Binary ClassificationConformal PredictionPrediction