paper-with-me

Papers

Evaluating Generative Models as Interactive Emergent Representations of Human-Like Collaborative Behavior

2026-05-05 · Shinas Shaji, Teena Chakkalayil Hassan, Sebastian Houben, Alex Mitrevski arxiv

Human-AI collaboration requires AI agents to understand human behavior for effective coordination. While advances in foundation models show promising capabilities in understanding and showing human-like behavior, their application in embodied collaborative settings needs further investigation. This work examines whether embodied foundation model agents exhibit emergent collaborative behaviors indicating underlying mental models of their collaborators, which is an important aspect of effective coordination. This paper develops a 2D collaborative game environment where large language model agents and humans complete color-matching tasks requiring coordination. We define five collaborative behaviors as indicators of emergent mental model representation: perspective-taking, collaborator-aware planning, introspection, theory of mind, and clarification. An automated behavior detection system using LLM-based judges identifies these behaviors, achieving fair to substantial agreement with human annotations. Results from the automated behavior detection system show that foundation models consistently exhibit emergent collaborative behaviors without being explicitly trained to do so. These behaviors occur at varying frequencies during collaboration stages, with distinct patterns across different LLMs. A user study was also conducted to evaluate human satisfaction and perceived collaboration effectiveness, with the results indicating positive collaboration experiences. Participants appreciated the agents' task focus, plan verbalization, and initiative, while suggesting improvements in response times and human-like interactions. This work provides an experimental framework for human-AI collaboration, empirical evidence of collaborative behaviors in embodied LLM agents, a validated behavioral analysis methodology, and an assessment of collaboration effectiveness.

📄 PDF Abstract BibTeX arXiv:2605.03855

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Generative UI: LLMs are Effective UI Generators

2026-02-24 · Yaniv Leviathan, Dani Valevski, Matan Kalman, Danny Lumen 외 arxiv

AI models excel at creating content, but typically render it with static, predefined interfaces. Specifically, the output of LLMs is often a markdown "wall of text". Generative UI is a long standing promise, where the mo…

Evaluating the Interpretability of Generative Models by Interactive Reconstruction

2021-02-02 · Andrew Slavin Ross, Nina Chen, Elisa Zhao Hang, Elena L. Glassman 외

For machine learning models to be most useful in numerous sociotechnical systems, many have argued that they must be human-interpretable. However, despite increasing interest in interpretability, there remains no firm co…

DisentanglementRepresentation Learning

Generative Agents: Interactive Simulacra of Human Behavior

2023-04-07 · Joon Sung Park, Joseph C. O'Brien, Carrie J. Cai, Meredith Ringel Morris 외

Believable proxies of human behavior can empower interactive applications ranging from immersive environments to rehearsal spaces for interpersonal communication to prototyping tools. In this paper, we introduce generati…

Language ModellingLarge Language Model

Emergent Communication in Interactive Sketch Question Answering

2023-09-21 · NeurIPS 2023 11

Vision-based emergent communication (EC) aims to learn to communicate through sketches and demystify the evolution of human communication. Ironically, previous works neglect multi-round interaction, which is indispensabl…

XferBench: a Data-Driven Benchmark for Emergent Language

2024-07-03 · Brendon Boldt, David Mortensen

In this paper, we introduce a benchmark for evaluating the overall quality of emergent languages using data-driven methods. Specifically, we interpret the notion of the "quality" of an emergent language as its similarity…