Agent policies from higher-order causal functions
We establish a correspondence between equivalence classes of agent-state policies for deterministic POMDPs and one-input process functions (the classical-deterministic limit of higher-order quantum operations). We use this correspondence to build a bridge between the agent-environment interaction in artificial intelligence, causal structure in the foundations of physics, and logic in computer science. We construct a *-autonomous category PF of types which supports an interpretation of one-step evaluation of policies, and multi-agent observation constraints, into cuts and monoidal products. In terms of types, we develop the correspondence further by identifying observation-independent decentralised POMDPs as the natural domain for the multi-input process functions used to model indefinite causality. We then prove a strict separation between general multi-input process function and definite-ordered process function performance on such dec-POMDPs, by finding an instance for which policies utilizing an indefinite causal structure can achieve greater finite-horizon rewards than policies which are restricted to a fixed background causal structure.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Causal Influence in Federated Edge Inference
In this paper, we consider a setting where heterogeneous agents with connectivity are performing inference using unlabeled streaming data. Observed data are only partially informative about the target variable of interes…
Crowd CountingDecision MakingPotential-Based Advice for Stochastic Policy Learning
This paper augments the reward received by a reinforcement learning agent with potential functions in order to help the agent learn (possibly stochastic) optimal policies. We show that a potential-based reward shaping sc…
Q-LearningReinforcement LearningTiered Reward: Designing Rewards for Specification and Fast Learning of Desired Behavior
Reinforcement-learning agents seek to maximize a reward signal through environmental interactions. As humans, our job in the learning process is to design reward functions to express desired behavior and enable the agent…
Deep Reinforcement Learningreinforcement-learningReinforcement LearningLearning Multi-Level Hierarchies with Hindsight
Hierarchical agents have the potential to solve sequential decision making tasks with greater sample efficiency than their non-hierarchical counterparts because hierarchical agents can break down tasks into sets of subta…
Decision MakingHierarchical Reinforcement LearningReinforcement LearningSequential Decision MakingQ-Cogni: An Integrated Causal Reinforcement Learning Framework
We present Q-Cogni, an algorithmically integrated causal reinforcement learning framework that redesigns Q-Learning with an autonomous causal structure discovery method to improve the learning process with causal inferen…
Causal InferenceDecision MakingQ-Learningreinforcement-learning+2