paper-with-me

Papers

Disentangling Exploration from Exploitation

2024-04-29 · Alessandro Lizzeri, Eran Shmaya, Leeat Yariv

Starting from Robbins (1952), the literature on experimentation via multi-armed bandits has wed exploration and exploitation. Nonetheless, in many applications, agents' exploration and exploitation need not be intertwined: a policymaker may assess new policies different than the status quo; an investor may evaluate projects outside her portfolio. We characterize the optimal experimentation policy when exploration and exploitation are disentangled in the case of Poisson bandits, allowing for general news structures. The optimal policy features complete learning asymptotically, exhibits lots of persistence, but cannot be identified by an index a la Gittins. Disentanglement is particularly valuable for intermediate parameter values.

📄 PDF Abstract BibTeX arXiv:2404.19116

Code (0)

등록된 구현이 없습니다.

Tasks

DisentanglementMulti-Armed Bandits

Similar Papers 제목 키워드 기반

DiPO: Disentangled Perplexity Policy Optimization for Fine-grained Exploration-Exploitation Trade-Off

2026-04-15 · Xiaofan Li, Ming Yang, Zhiyuan Ma, Shichao Ma 외 arxiv

Reinforcement Learning with Verifiable Rewards (RLVR) has catalyzed significant advances in the reasoning capabilities of Large Language Models (LLMs). However, effectively managing the exploration and exploitation trade…

Reinforcement LearningMathematical Reasoning

Disentangling Exploration of Large Language Models by Optimal Exploitation

2025-01-15 · Tim Grams, Patrick Betz, Christian Bartelt

Exploration is a crucial skill for self-improvement and open-ended problem-solving. However, it remains unclear if large language models can effectively explore the state-space within an unknown environment. This work is…

Prompt Engineering

MULEX: Disentangling Exploitation from Exploration in Deep RL

2019-07-01 · Lucas Beyer, Damien Vincent, Olivier Teboul, Sylvain Gelly 외

An agent learning through interactions should balance its action selection process between probing the environment to discover new rewards and using the information acquired in the past to adopt useful behaviour. This tr…

Deterministic Sequencing of Exploration and Exploitation for Reinforcement Learning

2022-09-12 · Piyush Gupta, Vaibhav Srivastava

We propose Deterministic Sequencing of Exploration and Exploitation (DSEE) algorithm with interleaving exploration and exploitation epochs for model-based RL problems that aim to simultaneously learn the system model, i.…

Efficient Explorationreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Dynamic Exploration-Exploitation Trade-Off in Active Learning Regression with Bayesian Hierarchical Modeling

2023-04-16 · Upala Junaida Islam, Kamran Paynabar, George Runger, Ashif Sikandar Iquebal

Active learning provides a framework to adaptively query the most informative experiments towards learning an unknown black-box function. Various approaches of active learning have been proposed in the literature, howeve…

Active Learningregression