paper-with-me

홈 › Papers

Multi-Objective Coordination Graphs for the Expected Scalarised Returns with Generative Flow Models

2022-07-01 · Conor F. Hayes, Timothy Verstraeten, Diederik M. Roijers, Enda Howley, Patrick Mannion

Many real-world problems contain multiple objectives and agents, where a trade-off exists between objectives. Key to solving such problems is to exploit sparse dependency structures that exist between agents. For example, in wind farm control a trade-off exists between maximising power and minimising stress on the systems components. Dependencies between turbines arise due to the wake effect. We model such sparse dependencies between agents as a multi-objective coordination graph (MO-CoG). In multi-objective reinforcement learning a utility function is typically used to model a users preferences over objectives, which may be unknown a priori. In such settings a set of optimal policies must be computed. Which policies are optimal depends on which optimality criterion applies. If the utility function of a user is derived from multiple executions of a policy, the scalarised expected returns (SER) must be optimised. If the utility of a user is derived from a single execution of a policy, the expected scalarised returns (ESR) criterion must be optimised. For example, wind farms are subjected to constraints and regulations that must be adhered to at all times, therefore the ESR criterion must be optimised. For MO-CoGs, the state-of-the-art algorithms can only compute a set of optimal policies for the SER criterion, leaving the ESR criterion understudied. To compute a set of optimal polices under the ESR criterion, also known as the ESR set, distributions over the returns must be maintained. Therefore, to compute a set of optimal policies under the ESR criterion for MO-CoGs, we present a novel distributional multi-objective variable elimination (DMOVE) algorithm. We evaluate DMOVE in realistic wind farm simulations. Given the returns in real-world wind farm settings are continuous, we utilise a model known as real-NVP to learn the continuous return distributions to calculate the ESR set.

📄 PDF Abstract BibTeX arXiv:2207.00368

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-Objective Reinforcement Learning

Similar Papers 제목 키워드 기반

A utility-based analysis of equilibria in multi-objective normal form games

2020-01-17 · Roxana Rădulescu, Patrick Mannion, Yijie Zhang, Diederik M. Roijers 외

In multi-objective multi-agent systems (MOMAS), agents explicitly consider the possible tradeoffs between conflicting objective functions. We argue that compromises between competing objectives in MOMAS should be analyse…

Form

Expected Scalarised Returns Dominance: A New Solution Concept for Multi-Objective Decision Making

2021-06-02 · Conor F. Hayes, Timothy Verstraeten, Diederik M. Roijers, Enda Howley 외

In many real-world scenarios, the utility of a user is derived from the single execution of a policy. In this case, to apply multi-objective reinforcement learning, the expected utility of the returns must be optimised. …

Decision MakingMulti-Objective Reinforcement Learningreinforcement-learningReinforcement Learning+1

Multi-Objective Multi-Agent Decision Making: A Utility-based Analysis and Survey

2019-09-06 · Roxana Rădulescu, Patrick Mannion, Diederik M. Roijers, Ann Nowé

The majority of multi-agent system (MAS) implementations aim to optimise agents' policies with respect to a single objective, despite the fact that many real-world problem domains are inherently multi-objective in nature…

Decision Making

A Demonstration of Issues with Value-Based Multiobjective Reinforcement Learning Under Stochastic State Transitions

2020-04-14 · Peter Vamplew, Cameron Foale, Richard Dazeley

We report a previously unidentified issue with model-free, value-based approaches to multiobjective reinforcement learning in the context of environments with stochastic state transitions. An example multiobjective Marko…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

An Empirical Investigation of Value-Based Multi-objective Reinforcement Learning for Stochastic Environments

2024-01-06 · Kewen Ding, Peter Vamplew, Cameron Foale, Richard Dazeley

One common approach to solve multi-objective reinforcement learning (MORL) problems is to extend conventional Q-learning by using vector Q-values in combination with a utility function. However issues can arise with this…

Multi-Objective Reinforcement LearningQ-Learning