paper-with-me

Papers

COMMA: A Communicative Multimodal Multi-Agent Benchmark

2024-10-10 · Timothy Ossowski, Jixuan Chen, Danyal Maqbool, Zefan Cai, Tyler Bradshaw, Junjie Hu

The rapid advances of multi-modal agents built on large foundation models have largely overlooked their potential for language-based communication between agents in collaborative tasks. This oversight presents a critical gap in understanding their effectiveness in real-world deployments, particularly when communicating with humans. Existing agentic benchmarks fail to address key aspects of inter-agent communication and collaboration, particularly in scenarios where agents have unequal access to information and must work together to achieve tasks beyond the scope of individual capabilities. To fill this gap, we introduce a novel benchmark designed to evaluate the collaborative performance of multimodal multi-agent systems through language communication. Our benchmark features a variety of scenarios, providing a comprehensive evaluation across four key categories of agentic capability in a communicative collaboration setting. By testing both agent-agent and agent-human collaborations using open-source and closed-source models, our findings reveal surprising weaknesses in state-of-the-art models, including proprietary models like GPT-4o. These models struggle to outperform even a simple random agent baseline in agent-agent collaboration and only surpass the random baseline when a human is involved.

📄 PDF Abstract BibTeX arXiv:2410.07553

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MUG: Interactive Multimodal Grounding on User Interfaces

2022-09-29 · Tao Li, Gang Li, Jingjie Zheng, Purple Wang 외

We present MUG, a novel interactive task for multimodal grounding where a user and an agent work collaboratively on an interface screen. Prior works modeled multimodal UI grounding in one round: the user gives a command …

VideoAgent: Personalized Synthesis of Scientific Videos

2025-09-14 · Xiao Liang, Bangxin Li, Zixuan Chen, Hanyue Zheng 외 arxiv

The technical complexity of research papers often limits their reach, necessitating more accessible formats like scientific videos to disseminate key insights through engaging narration. However, existing automated metho…

Voice-Interactive Surgical Agent for Multimodal Patient Data Control

2025-11-10 · Hyeryun Park, Byung Mo Gu, Jun Hee Lee, Byeong Hyeon Choi 외 arxiv

In robotic surgery, surgeons fully engage their hands and visual attention in procedures, making it difficult to access and manipulate multimodal patient data without interrupting the workflow. To overcome this problem, …

Interpreting and Generating Gestures with Embodied Human Computer Interactions

2020-09-18 · ACM IVA Workshop GENEA 2020 10 · Anonymous

In this paper, we discuss the role that gesture plays for an embodied intelligent virtual agent (IVA) in the context of multimodal task-oriented dialogues with a human. We have developed a simulation platform, VoxWorl…

Gesture Generation

CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society

2023-03-31 · NeurIPS 2023 11 · Guohao Li, Hasan Abed Al Kader Hammoud, Hani Itani, Dmitrii Khizbullin 외

The rapid advancement of chat-based language models has led to remarkable progress in complex task-solving. However, their success heavily relies on human input to guide the conversation, which can be challenging and tim…

Instruction FollowingLanguage ModelingLanguage ModellingLarge Language Model