paper-with-me

Papers

Towards Open-World Grasping with Large Vision-Language Models

2024-06-26 · Georgios Tziafas, Hamidreza Kasaei

The ability to grasp objects in-the-wild from open-ended language instructions constitutes a fundamental challenge in robotics. An open-world grasping system should be able to combine high-level contextual with low-level physical-geometric reasoning in order to be applicable in arbitrary scenarios. Recent works exploit the web-scale knowledge inherent in large language models (LLMs) to plan and reason in robotic context, but rely on external vision and action models to ground such knowledge into the environment and parameterize actuation. This setup suffers from two major bottlenecks: a) the LLM's reasoning capacity is constrained by the quality of visual grounding, and b) LLMs do not contain low-level spatial understanding of the world, which is essential for grasping in contact-rich scenarios. In this work we demonstrate that modern vision-language models (VLMs) are capable of tackling such limitations, as they are implicitly grounded and can jointly reason about semantics and geometry. We propose OWG, an open-world grasping pipeline that combines VLMs with segmentation and grasp synthesis models to unlock grounded world understanding in three stages: open-ended referring segmentation, grounded grasp planning and grasp ranking via contact reasoning, all of which can be applied zero-shot via suitable visual prompting mechanisms. We conduct extensive evaluation in cluttered indoor scene datasets to showcase OWG's robustness in grounding from open-ended language, as well as open-world robotic grasping experiments in both simulation and hardware that demonstrate superior performance compared to previous supervised and zero-shot LLM-based methods. Project material is available at https://gtziafas.github.io/OWG_project/ .

📄 PDF Abstract BibTeX arXiv:2406.18722

Code (0)

등록된 구현이 없습니다.

Tasks

Robotic GraspingVisual GroundingVisual Prompting

Similar Papers 제목 키워드 기반

GLOVER: Generalizable Open-Vocabulary Affordance Reasoning for Task-Oriented Grasping

2024-11-19 · Teli Ma, Zifan Wang, Jiaming Zhou, Mengmeng Wang 외

Inferring affordable (i.e., graspable) parts of arbitrary objects based on human specifications is essential for robots advancing toward open-vocabulary manipulation. Current grasp planners, however, are hindered by limi…

Common Sense ReasoningHuman-Object Interaction DetectionPose EstimationWorld Knowledge

OK-Robot: What Really Matters in Integrating Open-Knowledge Models for Robotics

2024-01-22 · Peiqi Liu, Yaswanth Orru, Jay Vakil, Chris Paxton 외

Remarkable progress has been made in recent years in the fields of vision, language, and robotics. We now have vision models capable of recognizing objects based on language queries, navigation systems that can effective…

object-detectionObject Detection

RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General Grasping

2025-07-31 · Dongming Wu, Yanping Fu, Saike Huang, Yingfei Liu 외 arxiv

General robotic grasping systems require accurate object affordance perception in diverse open-world scenarios following human instructions. However, current studies suffer from the problem of lacking reasoning-based lar…

Robot ManipulationRobotic Grasping

Grasp-MPC: Closed-Loop Visual Grasping via Value-Guided Model Predictive Control

2025-09-07 · Jun Yamada, Adithyavairavan Murali, Ajay Mandlekar, Clemens Eppner 외 arxiv

Grasping of diverse objects in unstructured environments remains a significant challenge. Open-loop grasping methods, effective in controlled settings, struggle in cluttered environments. Grasp prediction errors and obje…

Collision Avoidance

MetaGraspNet_v0: A Large-Scale Benchmark Dataset for Vision-driven Robotic Grasping via Physics-based Metaverse Synthesis

2021-12-29 · Yuhao Chen, E. Zhixuan Zeng, Maximilian Gilles, Alexander Wong

There has been increasing interest in smart factories powered by robotics systems to tackle repetitive, laborious tasks. One impactful yet challenging task in robotics-powered smart factory applications is robotic graspi…

Objectobject-detectionObject DetectionRobotic Grasping+1