paper-with-me

홈 › Papers

Click to Grasp: Zero-Shot Precise Manipulation via Visual Diffusion Descriptors

2024-03-21 · Nikolaos Tsagkas, Jack Rome, Subramanian Ramamoorthy, Oisin Mac Aodha, Chris Xiaoxuan Lu

Precise manipulation that is generalizable across scenes and objects remains a persistent challenge in robotics. Current approaches for this task heavily depend on having a significant number of training instances to handle objects with pronounced visual and/or geometric part ambiguities. Our work explores the grounding of fine-grained part descriptors for precise manipulation in a zero-shot setting by utilizing web-trained text-to-image diffusion-based generative models. We tackle the problem by framing it as a dense semantic part correspondence task. Our model returns a gripper pose for manipulating a specific part, using as reference a user-defined click from a source image of a visually different instance of the same object. We require no manual grasping demonstrations as we leverage the intrinsic object geometry and features. Practical experiments in a real-world tabletop scenario validate the efficacy of our approach, demonstrating its potential for advancing semantic-aware robotics manipulation. Web page: https://tsagkas.github.io/click2grasp

📄 PDF Abstract BibTeX arXiv:2403.14526

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Mana: Dexterous Manipulation of Articulated Tools

2026-06-11 · Zhao-Heng Yin, Guanya Shi, Pieter Abbeel, C. Karen Liu arxiv

Articulated tool manipulation remains a major challenge in dexterous robotics due to the need to coordinate internal degrees of freedom and contact-rich interactions. While prior work has largely focused on rigid objects…

Reinforcement LearningMotion Planning

A Training-Free Framework for Precise Mobile Manipulation of Small Everyday Objects

2025-02-19 · Arjun Gupta, Rishik Sathua, Saurabh Gupta

Many everyday mobile manipulation tasks require precise interaction with small objects, such as grasping a knob to open a cabinet or pressing a light switch. In this paper, we develop Servoing with Vision Models (SVM), a…

Imitation LearningPoint Tracking

NeuralTouch: Neural Descriptors for Precise Sim-to-Real Tactile Robot Control

2025-10-23 · Yijiong Lin, Bowen Deng, Keju Pu, Chenghua Lu 외 arxiv

Grasping accuracy is a critical prerequisite for precise object manipulation, often requiring careful alignment between the robot hand and object. Neural Descriptor Fields (NDF) offer a promising vision-based method to g…

Reinforcement LearningPoint Clouds

ZeroDexGrasp: Zero-Shot Task-Oriented Dexterous Grasp Synthesis with Prompt-Based Multi-Stage Semantic Reasoning

2025-11-17 · Juntao Jian, Yi-Lin Wei, Chengjie Mou, Yuhao Lin 외 arxiv

Task-oriented dexterous grasping holds broad application prospects in robotic manipulation and human-object interaction. However, most existing methods still struggle to generalize across diverse objects and task instruc…

Robotic Grasping

VLAD-Grasp: Zero-shot Grasp Detection via Vision-Language Models

2025-11-08 · Manav Kulshrestha, S. Talha Bukhari, Damon Conover, Aniket Bera arxiv

Robotic grasping is a fundamental capability for enabling autonomous manipulation, with usually infinite solutions. State-of-the-art approaches for grasping rely on learning from large-scale datasets comprising expert an…

Zero-shot GeneralizationRobotic GraspingPoint Clouds