paper-with-me

Papers

Sequential Counterfactual Risk Minimization

2023-02-23 · Houssam Zenati, Eustache Diemert, Matthieu Martin, Julien Mairal, Pierre Gaillard

Counterfactual Risk Minimization (CRM) is a framework for dealing with the logged bandit feedback problem, where the goal is to improve a logging policy using offline data. In this paper, we explore the case where it is possible to deploy learned policies multiple times and acquire new data. We extend the CRM principle and its theory to this scenario, which we call "Sequential Counterfactual Risk Minimization (SCRM)." We introduce a novel counterfactual estimator and identify conditions that can improve the performance of CRM in terms of excess risk and regret rates, by using an analysis similar to restart strategies in accelerated optimization methods. We also provide an empirical evaluation of our method in both discrete and continuous action settings, and demonstrate the benefits of multiple deployments of CRM.

📄 PDF Abstract BibTeX arXiv:2302.12120

Code (1)

criteo-research/sequential-conterfactual-risk-minimization 공식 구현 jax

Tasks

counterfactual

Similar Papers 제목 키워드 기반

Adversarial Counterfactual Environment Model Learning

2023-09-21 · NeurIPS 2023 11

An accurate environment dynamics model is crucial for various downstream tasks, such as counterfactual prediction, off-policy evaluation, and offline reinforcement learning. Currently, these models were learned through …

Distributionally Robust Counterfactual Risk Minimization

2019-06-14 · Louis Faury, Ugo Tanielian, Flavian vasile, Elena Smirnova 외

This manuscript introduces the idea of using Distributionally Robust Optimization (DRO) for the Counterfactual Risk Minimization (CRM) problem. Tapping into a rich existing literature, we show that DRO is a principled to…

counterfactualDecision Making

Bayesian Counterfactual Risk Minimization

2018-06-29 · Ben London, Ted Sandler

We present a Bayesian view of counterfactual risk minimization (CRM) for offline learning from logged bandit feedback. Using PAC-Bayesian analysis, we derive a new generalization bound for the truncated inverse propensit…

counterfactual

Stochastic Regret Minimization in Extensive-Form Games

2020-02-19 · ICML 2020 1 · Gabriele Farina, Christian Kroer, Tuomas Sandholm

Monte-Carlo counterfactual regret minimization (MCCFR) is the state-of-the-art algorithm for solving sequential games that are too large for full tree traversals. It works by using gradient estimates that can be computed…

counterfactualForm

Counterfactual Risk Minimization: Learning from Logged Bandit Feedback

2015-02-09 · Adith Swaminathan, Thorsten Joachims

We develop a learning principle and an efficient algorithm for batch learning from logged bandit feedback. This learning setting is ubiquitous in online systems (e.g., ad placement, web search, recommendation), where an …

counterfactualMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATION