paper-with-me

홈 › Papers

Negotiable Reinforcement Learning for Pareto Optimal Sequential Decision-Making

2018-12-01 · NeurIPS 2018 12 · Nishant Desai, Andrew Critch, Stuart J. Russell

It is commonly believed that an agent making decisions on behalf of two or more principals who have different utility functions should adopt a Pareto optimal policy, i.e. a policy that cannot be improved upon for one principal without making sacrifices for another. Harsanyi's theorem shows that when the principals have a common prior on the outcome distributions of all policies, a Pareto optimal policy for the agent is one that maximizes a fixed, weighted linear combination of the principals’ utilities. In this paper, we derive a more precise generalization for the sequential decision setting in the case of principals with different priors on the dynamics of the environment. We refer to this generalization as the Negotiable Reinforcement Learning (NRL) framework. In this more general case, the relative weight given to each principal’s utility should evolve over time according to how well the agent’s observations conform with that principal’s prior. To gain insight into the dynamics of this new framework, we implement a simple NRL agent and empirically examine its behavior in a simple environment.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Decision Makingreinforcement-learningReinforcement LearningReinforcement Learning (RL)Sequential Decision Making

Similar Papers 제목 키워드 기반

Toward negotiable reinforcement learning: shifting priorities in Pareto optimal sequential decision-making

2017-01-05 · Andrew Critch

Existing multi-objective reinforcement learning (MORL) algorithms do not account for objectives that arise from players with differing beliefs. Concretely, consider two players with different beliefs and utility function…

Decision MakingMulti-Objective Reinforcement Learningreinforcement-learningReinforcement Learning+2

Meta-Learning for Multi-objective Reinforcement Learning

2018-11-08 · Xi Chen, Ali Ghadirzadeh, Mårten Björkman, Patric Jensfelt

Multi-objective reinforcement learning (MORL) is the generalization of standard reinforcement learning (RL) approaches to solve sequential decision making problems that consist of several, possibly conflicting, objective…

Computational Efficiencycontinuous-controlContinuous ControlDecision Making+6

Pareto Inverse Reinforcement Learning for Diverse Expert Policy Generation

2024-08-22 · Woo Kyung Kim, Minjong Yoo, Honguk Woo

Data-driven offline reinforcement learning and imitation learning approaches have been gaining popularity in addressing sequential decision-making problems. Yet, these approaches rarely consider learning Pareto-optimal p…

Autonomous DrivingDecision MakingImitation Learningreinforcement-learning+2

Deterministic Pareto-Optimal Policy Synthesis for Multi-Objective Reinforcement Learning

2026-06-24 · Aniruddha Joshi, Niklas Lauffer, Sanjit Seshia arxiv

Real-world decision-making often requires balancing multiple conflicting objectives, a challenge that standard Reinforcement Learning (RL) frequently addresses by aggregating rewards into a single scalar signal. While ef…

Reinforcement Learning

Latent-Conditioned Policy Gradient for Multi-Objective Deep Reinforcement Learning

2023-03-15 · Takuya Kanazawa, Chetan Gupta

Sequential decision making in the real world often requires finding a good balance of conflicting objectives. In general, there exist a plethora of Pareto-optimal policies that embody different patterns of compromises be…

Decision MakingDeep Reinforcement LearningMulti-Objective Reinforcement Learningreinforcement-learning+3