paper-with-me

홈 › Papers

LLMs as Policy-Agnostic Teammates: A Case Study in Human Proxy Design for Heterogeneous Agent Teams

2025-10-07 · Aju Ani Justus, Chris Baber arxiv

A critical challenge in modelling Heterogeneous-Agent Teams is training agents to collaborate with teammates whose policies are inaccessible or non-stationary, such as humans. Traditional approaches rely on expensive human-in-the-loop data, which limits scalability. We propose using Large Language Models (LLMs) as policy-agnostic human proxies to generate synthetic data that mimics human decision-making. To evaluate this, we conduct three experiments in a grid-world capture game inspired by Stag Hunt, a game theory paradigm that balances risk and reward. In Experiment 1, we compare decisions from 30 human participants and 2 expert judges with outputs from LLaMA 3.1 and Mixtral 8x22B models. LLMs, prompted with game-state observations and reward structures, align more closely with experts than participants, demonstrating consistency in applying underlying decision criteria. Experiment 2 modifies prompts to induce risk-sensitive strategies (e.g. "be risk averse"). LLM outputs mirror human participants' variability, shifting between risk-averse and risk-seeking behaviours. Finally, Experiment 3 tests LLMs in a dynamic grid-world where the LLM agents generate movement actions. LLMs produce trajectories resembling human participants' paths. While LLMs cannot yet fully replicate human adaptability, their prompt-guided diversity offers a scalable foundation for simulating policy-agnostic teammates.

📄 PDF Abstract BibTeX arXiv:2510.06151

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Modelling the Dynamic Joint Policy of Teammates with Attention Multi-agent DDPG

2018-11-13 · Hangyu Mao, Zhengchao Zhang, Zhen Xiao, Zhibo Gong

Modelling and exploiting teammates' policies in cooperative multi-agent systems have long been an interest and also a big challenge for the reinforcement learning (RL) community. The interest lies in the fact that if the…

Reinforcement LearningReinforcement Learning (RL)

ProAgent: Building Proactive Cooperative Agents with Large Language Models

2023-08-22 · Ceyao Zhang, Kaijie Yang, Siyi Hu, ZiHao Wang 외

Building agents with adaptive behavior in cooperative tasks stands as a paramount goal in the realm of multi-agent systems. Current approaches to developing cooperative agents rely primarily on learning-based methods, wh…

Zero-Shot Coordination in Ad Hoc Teams with Generalized Policy Improvement and Difference Rewards

2025-10-17 · Rupal Nigam, Niket Parikh, Hamid Osooli, Mikihisa Yuasa 외 arxiv

Real-world multi-agent systems may require ad hoc teaming, where an agent must coordinate with other previously unseen teammates to solve a task in a zero-shot manner. Prior work often either selects a pretrained policy …

Generating Teammates for Training Robust Ad Hoc Teamwork Agents via Best-Response Diversity

2022-07-28 · Arrasy Rahman, Elliot Fosong, Ignacio Carlucho, Stefano V. Albrecht

Ad hoc teamwork (AHT) is the challenge of designing a robust learner agent that effectively collaborates with unknown teammates without prior coordination mechanisms. Early approaches address the AHT challenge by trainin…

Diversityvalid

A General Learning Framework for Open Ad Hoc Teamwork Using Graph-based Policy Learning

2022-10-11 · Arrasy Rahman, Ignacio Carlucho, Niklas Höpner, Stefano V. Albrecht

Open ad hoc teamwork is the problem of training a single agent to efficiently collaborate with an unknown group of teammates whose composition may change over time. A variable team composition creates challenges for the …

Graph Neural Network