Reward-Punishment Symmetric Universal Intelligence
Can an agent's intelligence level be negative? We extend the Legg-Hutter agent-environment framework to include punishments and argue for an affirmative answer to that question. We show that if the background encodings and Universal Turing Machine (UTM) admit certain Kolmogorov complexity symmetries, then the resulting Legg-Hutter intelligence measure is symmetric about the origin. In particular, this implies reward-ignoring agents have Legg-Hutter intelligence 0 according to such UTMs.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
The Effect of Punishment and Reward on Cooperation in a Prisoners' Dilemma Game
This work studies the effect of incentives (in the form of punishment and reward) on the equilibrium fraction of cooperators and defectors in an iterated n-person prisoners' dilemma game. With a finite population of play…
FormMediating Artificial Intelligence Developments through Negative and Positive Incentives
The field of Artificial Intelligence (AI) is going through a period of great expectations, introducing a certain level of anxiety in research, business and also policy. This anxiety is further energised by an AI race nar…
Hyperbolically-Discounted Reinforcement Learning on Reward-Punishment Framework
This paper proposes a new reinforcement learning with hyperbolic discounting. Combining a new temporal difference error with the hyperbolic discounting in recursive manner and reward-punishment framework, a new scheme to…
reinforcement-learningReinforcement LearningReinforcement Learning (RL)Adaptive Punishment for Cooperation in Mixed-Motive Games
Mixed-motive scenarios are ubiquitous in real-world multi-agent interactions, where self-interested agents often defect for immediate rewards, overlooking the potential of altruistic cooperation to improve long-term gain…
Regularized Reward-Punishment Reinforcement Learning
We propose KL-Coupled Policy Regularization (KCPR), a policy coordination framework for Reward-Punishment Reinforcement Learning (RPRL). Based on KCPR, we derive KL-Coupled Soft Optimality (KCSO) and develop its deep rea…
Reinforcement Learning