paper-with-me

홈 › Papers

OWMM-Agent: Open World Mobile Manipulation With Multi-modal Agentic Data Synthesis

2025-06-04 · Junting Chen, Haotian Liang, Lingxiao Du, Weiyun Wang, Mengkang Hu, Yao Mu, Wenhai Wang, Jifeng Dai, Ping Luo, Wenqi Shao, Lin Shao

The rapid progress of navigation, manipulation, and vision models has made mobile manipulators capable in many specialized tasks. However, the open-world mobile manipulation (OWMM) task remains a challenge due to the need for generalization to open-ended instructions and environments, as well as the systematic complexity to integrate high-level decision making with low-level robot control based on both global scene understanding and current agent state. To address this complexity, we propose a novel multi-modal agent architecture that maintains multi-view scene frames and agent states for decision-making and controls the robot by function calling. A second challenge is the hallucination from domain shift. To enhance the agent performance, we further introduce an agentic data synthesis pipeline for the OWMM task to adapt the VLM model to our task domain with instruction fine-tuning. We highlight our fine-tuned OWMM-VLM as the first dedicated foundation model for mobile manipulators with global scene understanding, robot state tracking, and multi-modal action generation in a unified model. Through experiments, we demonstrate that our model achieves SOTA performance compared to other foundation models including GPT-4o and strong zero-shot generalization in real world. The project page is at https://github.com/HHYHRHY/OWMM-Agent

📄 PDF Abstract BibTeX arXiv:2506.04217

Code (1)

hhyhrhy/owmm-agent 공식 구현

Tasks

Action GenerationDecision MakingHallucinationScene UnderstandingZero-shot Generalization

Similar Papers 제목 키워드 기반

HomeRobot Open Vocabulary Mobile Manipulation Challenge 2023 Participant Report (Team KuzHum)

2024-01-22 · Volodymyr Kuzma, Vladyslav Humennyy, Ruslan Partsey

We report an improvements to NeurIPS 2023 HomeRobot: Open Vocabulary Mobile Manipulation (OVMM) Challenge reinforcement learning baseline. More specifically, we propose more accurate semantic segmentation module, along w…

reinforcement-learningSemantic Segmentation

UniTeam: Open Vocabulary Mobile Manipulation Challenge

2023-12-14 · Andrew Melnik, Michael Büttner, Leon Harz, Lyon Brown 외

This report introduces our UniTeam agent - an improved baseline for the "HomeRobot: Open Vocabulary Mobile Manipulation" challenge. The challenge poses problems of navigation in unfamiliar environments, manipulation of n…

Object

Adaptive Mobile Manipulation for Articulated Objects In the Open World

2024-01-25 · Haoyu Xiong, Russell Mendonca, Kenneth Shaw, Deepak Pathak

Deploying robots in open-ended unstructured environments such as homes has been a long-standing research problem. However, robots are often studied only in closed-off lab settings, and prior mobile manipulation work is r…

EMMA: Scaling Mobile Manipulation via Egocentric Human Data

2025-09-04 · Lawrence Y. Zhu, Pranav Kuppili, Ryan Punamiya, Patcharapong Aphiwetsa 외 arxiv

Scaling mobile manipulation imitation learning is bottlenecked by expensive mobile robot teleoperation. We present Egocentric Mobile MAnipulation (EMMA), an end-to-end framework training mobile manipulation policies from…

Language-Conditioned Open-Vocabulary Mobile Manipulation with Pretrained Models

2025-07-23 · Shen Tan, Dong Zhou, Xiangyu Shao, Junqiao Wang 외 arxiv

Open-vocabulary mobile manipulation (OVMM) that involves the handling of novel and unseen objects across different workspaces remains a significant challenge for real-world robotic applications. In this paper, we propose…

Zero-shot GeneralizationMulti-Task Learning