Information-based learning by agents in unbounded state spaces
The idea that animals might use information-driven planning to explore an unknown environment and build an internal model of it has been proposed for quite some time. Recent work has demonstrated that agents using this principle can efficiently learn models of probabilistic environments with discrete, bounded state spaces. However, animals and robots are commonly confronted with unbounded environments. To address this more challenging situation, we study information-based learning strategies of agents in unbounded state spaces using non-parametric Bayesian models. Specifically, we demonstrate that the Chinese Restaurant Process (CRP) model is able to solve this problem and that an Empirical Bayes version is able to efficiently explore bounded and unbounded worlds by relying on little prior information.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Gradient-free Online Learning in Continuous Games with Delayed Rewards
Motivated by applications to online advertising and recommender systems, we consider a game-theoretic model with delayed rewards and asynchronous, payoff-based feedback. In contrast to previous work on delayed multi-arme…
Multi-Armed BanditsRecommendation SystemsBeyond Unbounded Beliefs: How Preferences and Information Interplay in Social Learning
When does society eventually learn the truth, or take the correct action, via observational learning? In a general model of sequential learning over social networks, we identify a simple condition for learning dubbed exc…
Concentration in unbounded metric spaces and algorithmic stability
We prove an extension of McDiarmid's inequality for metric spaces with unbounded diameter. To this end, we introduce the notion of the {\em subgaussian diameter}, which is a distribution-dependent refinement of the metri…
On Existence of Berk-Nash Equilibria in Misspecified Markov Decision Processes with Infinite Spaces
Model misspecification is a critical issue in many areas of theoretical and empirical economics. In the specific context of misspecified Markov Decision Processes, Esponda and Pouzo (2021) defined the notion of Berk-Nash…
Multiclass Transductive Online Learning
We consider the problem of multiclass transductive online learning when the number of labels can be unbounded. Previous works by Ben-David et al. [1997] and Hanneke et al. [2023b] only consider the case of binary and fin…