paper-with-me

홈 › Papers

Active Learning with Logged Data

2018-02-25 · ICML 2018 7 · Songbai Yan, Kamalika Chaudhuri, Tara Javidi

We consider active learning with logged data, where labeled examples are drawn conditioned on a predetermined logging policy, and the goal is to learn a classifier on the entire population, not just conditioned on the logging policy. Prior work addresses this problem either when only logged data is available, or purely in a controlled random experimentation setting where the logged data is ignored. In this work, we combine both approaches to provide an algorithm that uses logged data to bootstrap and inform experimentation, thus achieving the best of both worlds. Our work is inspired by a connection between controlled random experimentation and active learning, and modifies existing disagreement-based active learning algorithms to exploit logged data.

📄 PDF Abstract BibTeX arXiv:1802.09069

Code (0)

등록된 구현이 없습니다.

Tasks

Active Learning

Similar Papers 제목 키워드 기반

Active Offline Policy Selection

2021-06-18 · NeurIPS 2021 12 · Ksenia Konyushkova, Yutian Chen, Tom Le Paine, Caglar Gulcehre 외

This paper addresses the problem of policy selection in domains with abundant logged data, but with a restricted interaction budget. Solving this problem would enable safe evaluation and deployment of offline reinforceme…

Bayesian OptimizationOff-policy evaluation

A General Offline Reinforcement Learning Framework for Interactive Recommendation

2023-10-01 · Teng Xiao, Donglin Wang

This paper studies the problem of learning interactive recommender systems from logged feedbacks without any exploration in online environments. We address the problem by proposing a general offline reinforcement learnin…

Interactive RecommendationRecommendation Systemsreinforcement-learningReinforcement Learning

Off-Policy Evaluation and Learning from Logged Bandit Feedback: Error Reduction via Surrogate Policy

2018-08-01 · ICLR 2019 5 · Yuan Xie, Boyi Liu, Qiang Liu, Zhaoran Wang 외

When learning from a batch of logged bandit feedback, the discrepancy between the policy to be learned and the off-policy training data imposes statistical and computational challenges. Unlike classical supervised learni…

Multi-Label ClassificationMUlTI-LABEL-ClASSIFICATIONOff-policy evaluationRecommendation Systems

Demonstration of interactive teaching for end-to-end dialog control with hybrid code networks

2017-08-01 · WS 2017 8 · Jason D. Williams, Lars Liden

This is a demonstration of interactive teaching for practical end-to-end dialog systems driven by a recurrent neural network. In this approach, a developer teaches the network by interacting with the system and providing…

Dialog LearningEntity Extraction using GANIntent Detection

Model Inversion Networks for Model-Based Optimization

2019-12-31 · NeurIPS 2020 12 · Aviral Kumar, Sergey Levine

In this work, we aim to solve data-driven optimization problems, where the goal is to find an input that maximizes an unknown score function given access to a dataset of inputs with corresponding scores. When the inputs …

Bayesian Optimizationmodelvalid