paper-with-me

홈 › Papers

Bounding Boxes as Goals: Language-Conditioned Grasping via Neuro-Symbolic Planning

2026-06-11 · Allison Andreyev, Landon Eum, Nestor Tiglao, Romel Gomez arxiv

For robotics to be effectively integrated into household or industrial environments, machines must adapt to natural-language prompts in real time. Although Vision-Language Models (VLMs) have enabled zero-shot generalization in robot task and motion planning (TAMP), current state-of-the-art approaches often remain computationally "heavyweight" or require extensive training on thousands of demonstrations. We present GRASP (Grounded Reasoning and Symbolic Planning), a framework designed as a step toward open-vocabulary tabletop manipulation. Our approach leverages a pretrained VLM to translate natural-language queries into neuro-symbolic goal states, grounded in the physical world via a bounding-box detection pipeline. Unlike methods that rely on fixed color lists or hard-coded coordinates, GRASP enables robots to interpret abstract spatial concepts such as "top shelf" and execute tasks without additional fine-tuning. We achieve 73.3% overall success across 90 real-robot trials at three difficulty levels, requiring no task-specific training.

📄 PDF Abstract BibTeX arXiv:2606.12910

Code (0)

등록된 구현이 없습니다.

Tasks

Zero-shot GeneralizationMotion Planning

Similar Papers 제목 키워드 기반

RealVLG-R1: A Large-Scale Real-World Visual-Language Grounding Benchmark for Robotic Perception and Manipulation

2026-03-16 · Linfei Li, Lin Zhang, Ying Shen arxiv

Visual-language grounding aims to establish semantic correspondences between natural language and visual entities, enabling models to accurately identify and localize target objects based on textual instructions. Existin…

Robotic Grasping

Latent Particle World Models: Self-supervised Object-centric Stochastic Dynamics Modeling

2026-03-04 · Tal Daniel, Carl Qi, Dan Haramati, Amir Zadeh 외 arxiv

We introduce Latent Particle World Model (LPWM), a self-supervised object-centric world model scaled to real-world multi-object datasets and applicable in decision-making. LPWM autonomously discovers keypoints, bounding …

Diffusing More Objects for Semi-Supervised Domain Adaptation with Less Labeling

2023-12-19 · Leander van den Heuvel, Gertjan Burghouts, David W. Zhang, Gwenn Englebienne 외

For object detection, it is possible to view the prediction of bounding boxes as a reverse diffusion process. Using a diffusion model, the random bounding boxes are iteratively refined in a denoising step, conditioned on…

DenoisingDomain Adaptationobject-detectionObject Detection+1

Split, Merge, and Refine: Fitting Tight Bounding Boxes via Over-Segmentation and Iterative Search

2023-04-10 · Chanhyeok Park, Minhyuk Sung

Achieving tight bounding boxes of a shape while guaranteeing complete boundness is an essential task for efficient geometric operations and unsupervised semantic part detection. But previous methods fail to achieve both …

SegmentationSemantic Part DetectionSensitivity

NeuralLabeling: A versatile toolset for labeling vision datasets using Neural Radiance Fields

2023-09-21 · Floris Erich, Naoya Chiba, Yusuke Yoshiyasu, Noriaki Ando 외

We present NeuralLabeling, a labeling approach and toolset for annotating 3D scenes using either bounding boxes or meshes and generating segmentation masks, affordance maps, 2D bounding boxes, 3D bounding boxes, 6DOF obj…

Depth CompletionInstance SegmentationNeRFObject+2