paper-with-me

Papers

Policy Aggregation

2024-11-06 · Parand A. Alamdari, Soroush Ebadian, Ariel D. Procaccia

We consider the challenge of AI value alignment with multiple individuals that have different reward functions and optimal policies in an underlying Markov decision process. We formalize this problem as one of policy aggregation, where the goal is to identify a desirable collective policy. We argue that an approach informed by social choice theory is especially suitable. Our key insight is that social choice methods can be reinterpreted by identifying ordinal preferences with volumes of subsets of the state-action occupancy polytope. Building on this insight, we demonstrate that a variety of methods--including approval voting, Borda count, the proportional veto core, and quantile fairness--can be practically applied to policy aggregation.

📄 PDF Abstract BibTeX arXiv:2411.03651

Code (1)

praal/policy-aggregation 공식 구현

Tasks

Fairness

Similar Papers 제목 키워드 기반

Convergence of Value Aggregation for Imitation Learning

2018-01-22 · Ching-An Cheng, Byron Boots

Value aggregation is a general framework for solving imitation learning problems. Based on the idea of data aggregation, it generates a policy sequence by iteratively interleaving policy optimization and evaluation in an…

Imitation Learning

Feature-Based Aggregation and Deep Reinforcement Learning: A Survey and Some New Implementations

2018-04-12 · Dimitri P. Bertsekas

In this paper we discuss policy iteration methods for approximate solution of a finite-state discounted Markov decision problem, with a focus on feature-based aggregation methods and their connection with deep reinforcem…

Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

PolicyGNN: Aggregation Optimization for Graph Neural Networks

2020-02-01

Graph data are pervasive in many real-world applications. Recently, increasing attention has been paid on graph neural networks (GNNs), which aim to model the local graph structures and capture the hierarchical patterns …

Deep Reinforcement LearningReinforcement Learning (RL)

Policy-GNN: Aggregation Optimization for Graph Neural Networks

2020-06-26 · Kwei-Herng Lai, Daochen Zha, Kaixiong Zhou, Xia Hu

Graph data are pervasive in many real-world applications. Recently, increasing attention has been paid on graph neural networks (GNNs), which aim to model the local graph structures and capture the hierarchical patterns …

Deep Reinforcement LearningNode ClassificationReinforcement Learning (RL)

Biased Aggregation, Rollout, and Enhanced Policy Improvement for Reinforcement Learning

2019-10-06 · Dimitri Bertsekas

We propose a new aggregation framework for approximate dynamic programming, which provides a connection with rollout algorithms, approximate policy iteration, and other single and multistep lookahead methods. The central…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)