paper-with-me

홈 › Papers

GROW$^2$: Grounding Which and Where for Robot Tool Use

2026-06-29 · Yuhong Deng, Yuyao Liu, David Hsu arxiv

Can the robot use a plate to cut a cake if no knife is available? Tool use greatly expands robot capabilities, but to use tools creatively beyond their intended functions, the robot faces the challenge of $\textit{open-world affordance grounding}$: select an open-category object to act as a tool and localize its specific region of action. To this end, we introduce GROW$^2$ (GROunding Which and Where), which leverages object parts as a natural abstraction to split the grounding process hierarchically into semantic and geometric levels, thus bypassing the need for data-heavy, end-to-end training. Semantically, GROW$^2$ harnesses the commonsense reasoning of Vision-Language Models (VLMs) to parse a natural-language task instruction, select a suitable object as the tool, and identify task-relevant parts on the tool and the target object. Geometrically, vision foundation models then ground the selected parts into precise 3D regions from a single RGB-D image. Experiments on established benchmarks show that GROW$^2$ outperforms state-of-the-art baselines on affordance prediction benchmarks. Further, it achieves zero-shot generalization over open-category objects and outperforms baselines in both simulated and real-world robot tool use experiments.

📄 PDF Abstract BibTeX arXiv:2606.30632

Code (0)

등록된 구현이 없습니다.

Tasks

Zero-shot Generalization

Similar Papers 제목 키워드 기반

DSM: Building A Diverse Semantic Map for 3D Visual Grounding

2025-04-11 · Qinghongbing Xie, Zijian Liang, Long Zeng

In recent years, with the growing research and application of multimodal large language models (VLMs) in robotics, there has been an increasing trend of utilizing VLMs for robotic scene understanding tasks. Existing appr…

3D visual groundingScene UnderstandingSemantic SegmentationVisual Grounding

MedScope: Incentivizing "Think with Videos" for Clinical Reasoning via Coarse-to-Fine Tool Calling

2026-02-11 · Wenjie Li, Yujie Zhang, Haoran Sun, Xingqi He 외 arxiv

Long-form clinical videos are central to visual evidence-based decision-making, with growing importance for applications such as surgical robotics and related settings. However, current multimodal large language models t…

Gotta Grow Fast: Design and Benchmarking of a Tip Mount for High-Speed Vine Robots

2026-06-04 · Antonio Alvarez Valdivia, Robert Reeve, Ankush Dhawan, Ciera McFarland 외 arxiv

Soft, growing vine robots extend through tip eversion, a mechanism that enables navigation through cluttered environments. However, integrating cameras and other sensors at the tip is uniquely challenging because the mat…

Symbol, Conversational, and Societal Grounding with a Toy Robot

2017-09-29 · Casey Kennington, Sarah Plane

Essential to meaningful interaction is grounding at the symbolic, conversational, and societal levels. We present ongoing work with Anki's Cozmo toy robot as a research platform where we leverage the recent words-as-clas…

Creative Robot Tool Use by Counterfactual Reasoning

2026-05-06 · M. Tuluhan Akbulut, Varun Satheesh, Ahmed Jaafar, Alper Ahmetoglu 외 arxiv

We propose a causal reasoning framework for creative robot tool use where a suitable tool for a task is correctly identified for use beyond its primary objectives. The proposed framework first discovers the causal relati…