Value of Information and Reward Specification in Active Inference and POMDPs
Expected free energy (EFE) is a central quantity in active inference which has recently gained popularity due to its intuitive decomposition of the expected value of control into a pragmatic and an epistemic component. While numerous conjectures have been made to justify EFE as a decision making objective function, the most widely accepted is still its intuitiveness and resemblance to variational free energy in approximate Bayesian inference. In this work, we take a bottom up approach and ask: taking EFE as given, what's the resulting agent's optimality gap compared with a reward-driven reinforcement learning (RL) agent, which is well understood? By casting EFE under a particular class of belief MDP and using analysis tools from RL theory, we show that EFE approximates the Bayes optimal RL policy via information value. We discuss the implications for objective specification of active inference agents.
Code (0)
등록된 구현이 없습니다.
Tasks
Bayesian InferenceDecision MakingReinforcement Learning (RL)Similar Papers 제목 키워드 기반
Expectation Alignment: Handling Reward Misspecification in the Presence of Expectation Mismatch
Detecting and handling misspecified objectives, such as reward functions, has been widely recognized as one of the central challenges within the domain of Artificial Intelligence (AI) safety research. However, even with …
AI AgentInferring and Conveying Intentionality: Beyond Numerical Rewards to Logical Intentions
Shared intentionality is a critical component in developing conscious AI agents capable of collaboration, self-reflection, deliberation, and reasoning. We formulate inference of shared intentionality as an inverse reinfo…
reinforcement-learningReinforcement LearningReinforcement Learning (RL)Perceptual Values from Observation
Imitation by observation is an approach for learning from expert demonstrations that lack action information, such as videos. Recent approaches to this problem can be placed into two broad categories: training dynamics m…
reinforcement-learningReinforcement LearningReinforcement Learning (RL)Expressive Reward Synthesis with the Runtime Monitoring Language
A key challenge in reinforcement learning (RL) is reward (mis)specification, whereby imprecisely defined reward functions can result in unintended, possibly harmful, behaviours. Indeed, reward functions in RL are typical…
Reinforcement LearningSynthesizing Skeletons for Reactive Systems
We present an analysis technique for temporal specifications of reactive systems that identifies, on the level of individual system outputs over time, which parts of the implementation are determined by the specification…