paper-with-me

홈 › Papers

Reward-Punishment Symmetric Universal Intelligence

2021-10-06 · Samuel Allen Alexander, Marcus Hutter

Can an agent's intelligence level be negative? We extend the Legg-Hutter agent-environment framework to include punishments and argue for an affirmative answer to that question. We show that if the background encodings and Universal Turing Machine (UTM) admit certain Kolmogorov complexity symmetries, then the resulting Legg-Hutter intelligence measure is symmetric about the origin. In particular, this implies reward-ignoring agents have Legg-Hutter intelligence 0 according to such UTMs.

📄 PDF Abstract BibTeX arXiv:2110.02450

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Effect of Punishment and Reward on Cooperation in a Prisoners' Dilemma Game

2023-09-01 · Alexander Kangas

This work studies the effect of incentives (in the form of punishment and reward) on the equilibrium fraction of cooperators and defectors in an iterated n-person prisoners' dilemma game. With a finite population of play…

Form

Mediating Artificial Intelligence Developments through Negative and Positive Incentives

2020-10-01 · The Anh Han, Luis Moniz Pereira, Tom Lenaerts, Francisco C. Santos

The field of Artificial Intelligence (AI) is going through a period of great expectations, introducing a certain level of anxiety in research, business and also policy. This anxiety is further energised by an AI race nar…

Hyperbolically-Discounted Reinforcement Learning on Reward-Punishment Framework

2021-06-03 · Taisuke Kobayashi

This paper proposes a new reinforcement learning with hyperbolic discounting. Combining a new temporal difference error with the hyperbolic discounting in recursive manner and reward-punishment framework, a new scheme to…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Adaptive Punishment for Cooperation in Mixed-Motive Games

2026-05-23 · Min Tang, Fanqi Kong, Linyuan Lü, Xue Feng arxiv

Mixed-motive scenarios are ubiquitous in real-world multi-agent interactions, where self-interested agents often defect for immediate rewards, overlooking the potential of altruistic cooperation to improve long-term gain…

Regularized Reward-Punishment Reinforcement Learning

2026-06-26 · Jiexin Wang, Eiji Uchibe arxiv

We propose KL-Coupled Policy Regularization (KCPR), a policy coordination framework for Reward-Punishment Reinforcement Learning (RPRL). Based on KCPR, we derive KL-Coupled Soft Optimality (KCSO) and develop its deep rea…

Reinforcement Learning