paper-with-me

홈 › Papers

Non-maximizing policies that fulfill multi-criterion aspirations in expectation

2024-08-08 · Simon Dima, Simon Fischer, Jobst Heitzig, Joss Oliver

In dynamic programming and reinforcement learning, the policy for the sequential decision making of an agent in a stochastic environment is usually determined by expressing the goal as a scalar reward function and seeking a policy that maximizes the expected total reward. However, many goals that humans care about naturally concern multiple aspects of the world, and it may not be obvious how to condense those into a single reward function. Furthermore, maximization suffers from specification gaming, where the obtained policy achieves a high expected total reward in an unintended way, often taking extreme or nonsensical actions. Here we consider finite acyclic Markov Decision Processes with multiple distinct evaluation metrics, which do not necessarily represent quantities that the user wants to be maximized. We assume the task of the agent is to ensure that the vector of expected totals of the evaluation metrics falls into some given convex set, called the aspiration set. Our algorithm guarantees that this task is fulfilled by using simplices to approximate feasibility sets and propagate aspirations forward while ensuring they remain feasible. It has complexity linear in the number of possible state-action-successor triples and polynomial in the number of evaluation metrics. Moreover, the explicitly non-maximizing nature of the chosen policy and goals yields additional degrees of freedom, which can be used to apply heuristic safety criteria to the choice of actions. We discuss several such safety criteria that aim to steer the agent towards more conservative behavior.

📄 PDF Abstract BibTeX arXiv:2408.04385

Code (0)

등록된 구현이 없습니다.

Tasks

Sequential Decision Making

Similar Papers 제목 키워드 기반

MaxMI: A Maximal Mutual Information Criterion for Manipulation Concept Discovery

2024-07-21 · Pei Zhou, Yanchao Yang

We aim to discover manipulation concepts embedded in the unannotated demonstrations, which are recognized as key physical states. The discovered concepts can facilitate training manipulation policies and promote generali…

Would Friedman Burn your Tokens?

2023-06-29 · Aggelos Kiayias, Philip Lazos, Jan Christoph Schlegel

Cryptocurrencies come with a variety of tokenomic policies as well as aspirations of desirable monetary characteristics that have been described by proponents as 'sound money' or even 'ultra sound money.' These propositi…

Peer Networks and Malleability of Educational Aspirations

2022-09-17 · Michelle González Amador, Robin Cowan, Eleonora Nillesen

Continuing education beyond the compulsory years of schooling is one of the most important choices an adolescent has to make. Higher education is associated with a host of social and economic benefits both for the person…

Versatile Skill Control via Self-supervised Adversarial Imitation of Unlabeled Mixed Motions

2022-09-16 · Chenhao Li, Sebastian Blaes, Pavel Kolev, Marin Vlastelica 외

Learning diverse skills is one of the main challenges in robotics. To this end, imitation learning approaches have achieved impressive results. These methods require explicitly labeled datasets or assume consistent skill…

Imitation Learning

Hope, Aspirations, and the Impact of LLMs on Female Programming Learners in Afghanistan

2025-11-09 · Hamayoon Behmanush, Freshta Akhtari, Roghieh Nooripour, Ingmar Weber 외 arxiv

Designing impactful educational technologies in contexts of socio-political instability requires a nuanced understanding of educational aspirations. Currently, scalable metrics for measuring aspirations are limited. This…