paper-with-me

홈 › Papers

$\mathcal{P}^3$: Toward Versatile Embodied Agents

2025-08-09 · Shengli Zhou, Xiangchen Wang, Jinrui Zhang, Ruozai Tian, Rongtao Xu, Guanhua Chen, Feng Zheng arxiv

Embodied agents have shown promising generalization capabilities across diverse physical environments, making them essential for a wide range of real-world applications. However, building versatile embodied agents poses critical challenges due to three key issues: dynamic environment perception, open-ended tool usage, and complex multi-task planning. Most previous works rely solely on feedback from tool agents to perceive environmental changes and task status, which limits adaptability to real-time dynamics, causes error accumulation, and restricts tool flexibility. Furthermore, multi-task scheduling has received limited attention, primarily due to the inherent complexity of managing task dependencies and balancing competing priorities in dynamic and complex environments. To overcome these challenges, we introduce $\mathcal P^3$, a unified framework that integrates real-time perception and dynamic scheduling. Specifically, $\mathcal P^3$ enables agents to perceive task-relevant information actively from the environment, plug and utilize tools without feedback requirements, and plan multi-task execution by prioritizing urgent tasks and dynamically adjusting task order based on dependencies. Extensive real-world experiments show that our approach bridges the gap between benchmarks and practical deployment, delivering highly transferable, general-purpose embodied agents. Code and data are available at https://github.com/fz-zsl/P3.

📄 PDF Abstract BibTeX arXiv:2508.07033

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language Models

2022-12-08 · ICCV 2023 1 · Chan Hee Song, Jiaman Wu, Clayton Washington, Brian M. Sadler 외

This study focuses on using large language models (LLMs) as a planner for embodied agents that can follow natural language instructions to complete complex tasks in a visually-perceived environment. The high data cost an…

VLNVerse: A Benchmark for Vision-Language Navigation with Versatile, Embodied, Realistic Simulation and Evaluation

2025-12-22 · Sihao Lin, Zerui Li, Xunyi Zhao, Gengze Zhou 외 arxiv

Despite remarkable progress in Vision-Language Navigation (VLN), existing benchmarks remain confined to fixed, small-scale datasets with naive physical simulation. These shortcomings limit the insight that the benchmarks…

Vision-Language Navigation

LangSuitE: Planning, Controlling and Interacting with Large Language Models in Embodied Text Environments

2024-06-24 · Zixia Jia, Mengmeng Wang, Baichen Tong, Song-Chun Zhu 외

Recent advances in Large Language Models (LLMs) have shown inspiring achievements in constructing autonomous agents that rely on language descriptions as inputs. However, it remains unclear how well LLMs can function as …

World Knowledge

Bridging VLMs and Embodied Intelligence with Deliberate Practice Policy Optimization

2025-11-20 · Yi Zhang, Che Liu, Xiancong Ren, Hanchu Ni 외 arxiv

Developing a universal and versatile embodied intelligence system presents two primary challenges: the critical embodied data bottleneck, where real-world data is scarce and expensive, and the algorithmic inefficiency of…

Reinforcement Learning

Do We Really Need a Complex Agent System? Distill Embodied Agent into a Single Model

2024-04-06 · Zhonghan Zhao, Ke Ma, Wenhao Chai, Xuan Wang 외

With the power of large language models (LLMs), open-ended embodied agents can flexibly understand human instructions, generate interpretable guidance strategies, and output executable actions. Nowadays, Multi-modal Lang…

Knowledge Distillation