paper-with-me

Papers

PhysiAgent: An Embodied Agent Framework in Physical World

2025-09-29 · Zhihao Wang, Jianxiong Li, Jinliang Zheng, Wencong Zhang, Dongxiu Liu, Yinan Zheng, Haoyi Niu, Junzhi Yu, Xianyuan Zhan arxiv

Vision-Language-Action (VLA) models have achieved notable success but often struggle with limited generalizations. To address this, integrating generalized Vision-Language Models (VLMs) as assistants to VLAs has emerged as a popular solution. However, current approaches often combine these models in rigid, sequential structures: using VLMs primarily for high-level scene understanding and task planning, and VLAs merely as executors of lower-level actions, leading to ineffective collaboration and poor grounding challenges. In this paper, we propose an embodied agent framework, PhysiAgent, tailored to operate effectively in physical environments. By incorporating monitor, memory, self-reflection mechanisms, and lightweight off-the-shelf toolboxes, PhysiAgent offers an autonomous scaffolding framework to prompt VLMs to organize different components based on real-time proficiency feedback from VLAs to maximally exploit VLAs' capabilities. Experimental results demonstrate significant improvements in task-solving performance on complex real-world robotic tasks, showcasing effective self-regulation of VLMs, coherent tool collaboration, and adaptive evolution of the framework during execution. PhysiAgent makes practical and pioneering efforts to integrate VLMs and VLAs, effectively grounding embodied agent frameworks in real-world settings.

📄 PDF Abstract BibTeX arXiv:2509.24524

Code (0)

등록된 구현이 없습니다.

Tasks

Scene Understanding

Similar Papers 제목 키워드 기반

HoloAgent-0: A Unified Embodied Agent Framework with 3D Spatial Memory

2026-06-22 · Xiaolin Zhou, Liu Liu, Tingyang Xiao, Wei Feng 외 arxiv

LLM agents follow a practical execution loop in digital environments: they reason over structured states, invoke tools, inspect feedback, and revise actions. Extending this loop to physical robots is difficult because ph…

Embodied AI Agents: Modeling the World

2025-06-27 · Pascale Fung, Yoram Bachrach, Asli Celikyilmaz, Kamalika Chaudhuri 외

This paper describes our research on AI agents embodied in visual, virtual or physical forms, enabling them to interact with both users and their environments. These agents, which include virtual avatars, wearable device…

Human Agent Collaboration

Hy-Embodied-VLM-1.0: Efficient Physical-World Agents

2026-07-14 · Ziyi Wang, Xumin Yu, Yongming Rao, Yonggen Ling 외 arxiv

Building capable embodied agents requires not only multimodal perception and understanding, but also agentic capabilities for reasoning about actions, adapting to evolving situations, and interacting with the physical wo…

PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding

2025-01-27 · Wei Chow, Jiageng Mao, Boyi Li, Daniel Seita 외

Understanding the physical world is a fundamental challenge in embodied AI, critical for enabling agents to perform complex tasks and operate safely in real-world environments. While Vision-Language Models (VLMs) have sh…

BenchmarkingCommon Sense ReasoningScene UnderstandingTask Planning

Embodied Science: Closing the Discovery Loop with Agentic Embodied AI

2026-03-20 · Xiang Zhuang, Chenyi Zhou, Kehua Feng, Zhihui Zhu 외 arxiv

Artificial intelligence has demonstrated remarkable capability in predicting scientific properties, yet scientific discovery remains an inherently physical, long-horizon pursuit governed by experimental cycles. Most curr…