paper-with-me

홈 › Papers

Object Segmentation from Open-Vocabulary Manipulation Instructions Based on Optimal Transport Polygon Matching with Multimodal Foundation Models

2024-07-01 · Takayuki Nishimura, Katsuyuki Kuyo, Motonari Kambara, Komei Sugiura

We consider the task of generating segmentation masks for the target object from an object manipulation instruction, which allows users to give open vocabulary instructions to domestic service robots. Conventional segmentation generation approaches often fail to account for objects outside the camera's field of view and cases in which the order of vertices differs but still represents the same polygon, which leads to erroneous mask generation. In this study, we propose a novel method that generates segmentation masks from open vocabulary instructions. We implement a novel loss function using optimal transport to prevent significant loss where the order of vertices differs but still represents the same polygon. To evaluate our approach, we constructed a new dataset based on the REVERIE dataset and Matterport3D dataset. The results demonstrated the effectiveness of the proposed method compared with existing mask generation methods. Remarkably, our best model achieved a +16.32% improvement on the dataset compared with a representative polygon-based method.

📄 PDF Abstract BibTeX arXiv:2407.00985

Code (0)

등록된 구현이 없습니다.

Tasks

SegmentationSemantic Segmentation

Methods 이 논문이 사용한 방법론

Golden Queue Managers 설명 없음

Similar Papers 제목 키워드 기반

Future Success Prediction in Open-Vocabulary Object Manipulation Tasks Based on End-Effector Trajectories

2024-12-26 · Motonari Kambara, Komei Sugiura

This study addresses a task designed to predict the future success or failure of open-vocabulary object manipulation. In this task, the model is required to make predictions based on natural language instructions, egocen…

ObjectPrediction

KUDA: Keypoints to Unify Dynamics Learning and Visual Prompting for Open-Vocabulary Robotic Manipulation

2025-03-13 · Zixian Liu, Mingtong Zhang, Yunzhu Li

With the rapid advancement of large language models (LLMs) and vision-language models (VLMs), significant progress has been made in developing open-vocabulary robotic manipulation systems. However, many existing approach…

ObjectVisual Prompting

Rethinking Intermediate Representation for VLM-based Robot Manipulation

2025-11-24 · Weiliang Tang, Jialin Gao, Jia-Hui Pan, Gang Wang 외 arxiv

Vision-Language Model (VLM) is an important component to enable robust robot manipulation. Yet, using it to translate human instructions into an action-resolvable intermediate representation often needs a tradeoff betwee…

Robot ManipulationFew-Shot Learning

Language-Conditioned Open-Vocabulary Mobile Manipulation with Pretrained Models

2025-07-23 · Shen Tan, Dong Zhou, Xiangyu Shao, Junqiao Wang 외 arxiv

Open-vocabulary mobile manipulation (OVMM) that involves the handling of novel and unseen objects across different workspaces remains a significant challenge for real-world robotic applications. In this paper, we propose…

Zero-shot GeneralizationMulti-Task Learning

DM2RM: Dual-Mode Multimodal Ranking for Target Objects and Receptacles Based on Open-Vocabulary Instructions

2024-08-15 · Ryosuke Korekata, Kanta Kaneda, Shunya Nagashima, Yuto Imai 외

In this study, we aim to develop a domestic service robot (DSR) that, guided by open-vocabulary instructions, can carry everyday objects to the specified pieces of furniture. Few existing methods handle mobile manipulati…

Image RetrievalLanguage ModellingLarge Language ModelRetrieval