paper-with-me

홈 › Papers

EMAC+: Embodied Multimodal Agent for Collaborative Planning with VLM+LLM

2025-05-26 · Shuang Ao, Flora D. Salim, Simon Khan

Although LLMs demonstrate proficiency in several text-based reasoning and planning tasks, their implementation in robotics control is constrained by significant deficiencies: (1) LLM agents are designed to work mainly with textual inputs rather than visual conditions; (2) Current multimodal agents treat LLMs as static planners, which separates their reasoning from environment dynamics, resulting in actions that do not take domain-specific knowledge into account; and (3) LLMs are not designed to learn from visual interactions, which makes it harder for them to make better policies for specific domains. In this paper, we introduce EMAC+, an Embodied Multimodal Agent that collaboratively integrates LLM and VLM via a bidirectional training paradigm. Unlike existing methods, EMAC+ dynamically refines high-level textual plans generated by an LLM using real-time feedback from a VLM executing low-level visual control tasks. We address critical limitations of previous models by enabling the LLM to internalize visual environment dynamics directly through interactive experience, rather than relying solely on static symbolic mappings. Extensive experimental evaluations on ALFWorld and RT-1 benchmarks demonstrate that EMAC+ achieves superior task performance, robustness against noisy observations, and efficient learning. We also conduct thorough ablation studies and provide detailed analyses of success and failure cases.

📄 PDF Abstract BibTeX arXiv:2505.19905

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Advances and Innovations in the Multi-Agent Robotic System (MARS) Challenge

2026-01-26 · Li Kang, Heng Zhou, Xiufeng Song, Rui Li 외 arxiv

Recent advancements in multimodal large language models and vision-languageaction models have significantly driven progress in Embodied AI. As the field transitions toward more complex task scenarios, multi-agent system …

Collaborative Tree Search for Enhancing Embodied Multi-Agent Collaboration

2025-01-01 · CVPR 2025 1 · Lizheng Zu, Lin Lin, Song Fu, Na Zhao 외

Embodied agents based on large language models (LLMs) face significant challenges in collaborative tasks, requiring effective communication and reasonable division of labor to ensure efficient and correct task comple…

Octopus: Embodied Vision-Language Programmer from Environmental Feedback

2023-10-12 · Jingkang Yang, Yuhao Dong, Shuai Liu, Bo Li 외

Large vision-language models (VLMs) have achieved substantial progress in multimodal perception and reasoning. When integrated into an embodied agent, existing embodied VLM works either output detailed action sequences a…

BenchmarkingCode GenerationDecision MakingMinecraft

Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses

2026-03-28 · Xiao Li, Xiang Zheng, Yifeng Gao, Xinyu Xia 외 arxiv

Embodied Artificial Intelligence (Embodied AI) integrates perception, cognition, planning, and interaction into agents that operate in open-world, safety-critical environments. As these systems gain autonomy and enter do…

MECoBench: A Systematic Study of Multimodal Agent Collaboration in Embodied Environments

2026-06-30 · Qingyun Liu, Jiwen Zhang, Jingyi Hu, Siyuan Wang 외 arxiv

Recent multimodal large language models (MLLMs) have strong potential as embodied agents, but their ability to collaborate in visually grounded environments remains underexplored. To address this gap, we introduce MECoBe…