paper-with-me

Papers

ELIGN: Expectation Alignment as a Multi-Agent Intrinsic Reward

2022-10-09 · Zixian Ma, Rose Wang, Li Fei-Fei, Michael Bernstein, Ranjay Krishna

Modern multi-agent reinforcement learning frameworks rely on centralized training and reward shaping to perform well. However, centralized training and dense rewards are not readily available in the real world. Current multi-agent algorithms struggle to learn in the alternative setup of decentralized training or sparse rewards. To address these issues, we propose a self-supervised intrinsic reward ELIGN - expectation alignment - inspired by the self-organization principle in Zoology. Similar to how animals collaborate in a decentralized manner with those in their vicinity, agents trained with expectation alignment learn behaviors that match their neighbors' expectations. This allows the agents to learn collaborative behaviors without any external reward or centralized training. We demonstrate the efficacy of our approach across 6 tasks in the multi-agent particle and the complex Google Research football environments, comparing ELIGN to sparse and curiosity-based intrinsic rewards. When the number of agents increases, ELIGN scales well in all multi-agent tasks except for one where agents have different capabilities. We show that agent coordination improves through expectation alignment because agents learn to divide tasks amongst themselves, break coordination symmetries, and confuse adversaries. These results identify tasks where expectation alignment is a more useful strategy than curiosity-driven exploration for multi-agent coordination, enabling agents to do zero-shot coordination.

📄 PDF Abstract BibTeX arXiv:2210.04365

Code (1)

stanfordvl/alignment 공식 구현 pytorch

Tasks

Multi-agent Reinforcement Learning

Methods 이 논문이 사용한 방법론

((Reservation@Faqs))How do I cancel a reservation on Expedia? How do I cancel a reservation on Expedia? +1^888^829^0881° oR +1^888^829^0881 – Need to cancel your Expedia reservation quickly and without hassle? This step-by-step guide…
Six Ways To Communicate To Someone At Expedia Via Phone And Email's. To communicate or get human at Expedia, the quickest option is typically to call their customer service at +1-888-829-0881 or +1(805) 330 (4056). You can also use the live chat…
ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
1x1 Convolution A 1 x 1 Convolution is a convolution with some special properties in that it can be used for dimensionality reduction,…
Feedforward Network A Feedforward Network, or a Multilayer Perceptron (MLP), is a neural network with solely densely connected layers. This is the classic neural network architecture of the…
TTUR The Two Time-scale Update Rule (TTUR) is an update rule for generative adversarial networks trained with stochastic gradient descent. TTUR has an individual learning rate for…
Projection Discriminator A Projection Discriminator is a type of discriminator for generative adversarial networks. It is motivated by a probabilistic model in which the distribution of the…

Similar Papers 제목 키워드 기반

Elign: Equivariant Diffusion Model Alignment from Foundational Machine Learning Force Fields

2026-01-29 · Yunyang Li, Lin Huang, Luojia Xia, Wenhe Zhang 외 arxiv

Generative models for 3D molecular conformations must respect Euclidean symmetries and concentrate probability mass on thermodynamically favorable, mechanically stable structures. However, E(3)-equivariant diffusion mode…

Reinforcement Learning

Active Alignments of Lens Systems with Reinforcement Learning

2025-03-03 · Matthias Burkhardt, Tobias Schmähling, Michael Layh, Tobias Windisch

Aligning a lens system relative to an imager is a critical challenge in camera manufacturing. While optimal alignment can be mathematically computed under ideal conditions, real-world deviations caused by manufacturing t…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

The Shadow Self: Intrinsic Value Misalignment in Large Language Model Agents

2026-01-24 · Chen Chen, Kim Young Il, Yuan Yang, Wenhao Su 외 arxiv

Large language model (LLM) agents with extended autonomy unlock new capabilities, but also introduce heightened challenges for LLM safety. In particular, an LLM agent may pursue objectives that deviate from human values …

New Alignment Methods for Discriminative Book Summarization

2013-05-06 · David Bamman, Noah A. Smith

We consider the unsupervised alignment of the full text of a book with a human-written summary. This presents challenges not seen in other text alignment problems, including a disparity in length and, consequent to this,…

Book summarization

Reliably Re-Acting to Partner's Actions with the Social Intrinsic Motivation of Transfer Empowerment

2022-03-07 · Tessa van der Heiden, Herke van Hoof, Efstratios Gavves, Christoph Salge

We consider multi-agent reinforcement learning (MARL) for cooperative communication and coordination tasks. MARL agents can be brittle because they can overfit their training partners' policies. This overfitting can prod…

Multi-agent Reinforcement Learningreinforcement-learningReinforcement Learning (RL)