paper-with-me

Papers

A Joint Network for Grasp Detection Conditioned on Natural Language Commands

2021-04-01 · Yiye Chen, Ruinian Xu, Yunzhi Lin, Patricio A. Vela

We consider the task of grasping a target object based on a natural language command query. Previous work primarily focused on localizing the object given the query, which requires a separate grasp detection module to grasp it. The cascaded application of two pipelines incurs errors in overlapping multi-object cases due to ambiguity in the individual outputs. This work proposes a model named Command Grasping Network(CGNet) to directly output command satisficing grasps from RGB image and textual command inputs. A dataset with ground truth (image, command, grasps) tuple is generated based on the VMRD dataset to train the proposed network. Experimental results on the generated test set show that CGNet outperforms a cascaded object-retrieval and grasp detection baseline by a large margin. Three physical experiments demonstrate the functionality and performance of CGNet.

📄 PDF Abstract BibTeX arXiv:2104.00492

Code (0)

등록된 구현이 없습니다.

Tasks

ObjectRetrieval

Similar Papers 제목 키워드 기반

Speech2Grasp: Data-Efficient Transfer of Text-Conditioned Grasp Detection to Speech in Humanoid Robots

2026-07-29 · Hung Nguyen, Kim Nhat Minh Nguyen, Van Duc Vu, Van-Danh Le 외 arxiv

Humanoid robots increasingly require multi-modal understanding for natural interaction with humans. Despite the prominence of vision-language models, they generally assume textual rather than the more natural speech inpu…

Language-Guided Grasp Detection with Coarse-to-Fine Learning for Robotic Manipulation

2025-12-24 · Zebin Jiang, Tianle Jin, Xiangtong Yao, Alois Knoll 외 arxiv

Grasping is one of the most fundamental challenging capabilities in robotic manipulation, especially in unstructured, cluttered, and semantically diverse environments. Recent researches have increasingly explored languag…

Bounding Boxes as Goals: Language-Conditioned Grasping via Neuro-Symbolic Planning

2026-06-11 · Allison Andreyev, Landon Eum, Nestor Tiglao, Romel Gomez arxiv

For robotics to be effectively integrated into household or industrial environments, machines must adapt to natural-language prompts in real time. Although Vision-Language Models (VLMs) have enabled zero-shot generalizat…

Zero-shot GeneralizationMotion Planning

Language-driven Grasp Detection with Mask-guided Attention

2024-07-29 · Tuan Van Vo, Minh Nhat Vu, Baoru Huang, An Vuong 외

Grasp detection is an essential task in robotics with various industrial applications. However, traditional methods often struggle with occlusions and do not utilize language for grasping. Incorporating natural language …

Semantic Segmentation

Language-Driven 6-DoF Grasp Detection Using Negative Prompt Guidance

2024-07-18 · Toan Nguyen, Minh Nhat Vu, Baoru Huang, An Vuong 외

6-DoF grasp detection has been a fundamental and challenging problem in robotic vision. While previous works have focused on ensuring grasp stability, they often do not consider human intention conveyed through natural l…

Benchmarking