paper-with-me

홈 › Papers

Generalizable Coarse-to-Fine Robot Manipulation via Language-Aligned 3D Keypoints

2025-09-28 · Jianshu Hu, Lidi Wang, Shujia Li, Yunpeng Jiang, Xiao Li, Paul Weng, Yutong Ban arxiv

Hierarchical coarse-to-fine policy, where a coarse branch predicts a region of interest to guide a fine-grained action predictor, has demonstrated significant potential in robotic 3D manipulation tasks by especially enhancing sample efficiency and enabling more precise manipulation. However, even augmented with pre-trained models, these hierarchical policies still suffer from generalization issues. To enhance generalization to novel instructions and environment variations, we propose Coarse-to-fine Language-Aligned manipulation Policy (CLAP), a framework that integrates three key components: 1) task decomposition, 2) VLM fine-tuning for 3D keypoint prediction, and 3) 3D-aware representation. Through comprehensive experiments in simulation and on a real robot, we demonstrate its superior generalization capability. Specifically, on GemBench, a benchmark designed for evaluating generalization, our approach achieves a 12\% higher average success rate than the SOTA method while using only 1/5 of the training trajectories. In real-world experiments, our policy, trained on only 10 demonstrations, successfully generalizes to novel instructions and environments.

📄 PDF Abstract BibTeX arXiv:2509.23575

Code (0)

등록된 구현이 없습니다.

Tasks

Robot Manipulation

Similar Papers 제목 키워드 기반

VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation

2025-03-10 · CVPR 2025 1 · Hanzhi Chen, Boyang Sun, Anran Zhang, Marc Pollefeys 외

Future robots are envisioned as versatile systems capable of performing a variety of household tasks. The big question remains, how can we bridge the embodiment gap while minimizing physical robot learning, which fundame…

Affordance-Guided Coarse-to-Fine Exploration for Base Placement in Open-Vocabulary Mobile Manipulation

2025-11-09 · Tzu-Jung Lin, Jia-Fong Yeh, Hung-Ting Su, Chung-Yi Lin 외 arxiv

In open-vocabulary mobile manipulation (OVMM), task success often hinges on the selection of an appropriate base placement for the robot. Existing approaches typically navigate to proximity-based regions without consider…

Multimodal Reasoning

GEAR-VLA: Learning Geometry-Aware Action Representations for Generalizable Robotic Manipulation

2026-06-07 · Yuan Zhang, Shiqi Zhang, Yedong Shen, Shuai Dong 외 arxiv

Vision-Language-Action (VLA) models achieve strong benchmark performance but still struggle in real-world deployment with unseen objects, background shifts, and different robot embodiments. We argue that this stems from …

Action Understanding

Language-Guided Grasp Detection with Coarse-to-Fine Learning for Robotic Manipulation

2025-12-24 · Zebin Jiang, Tianle Jin, Xiangtong Yao, Alois Knoll 외 arxiv

Grasping is one of the most fundamental challenging capabilities in robotic manipulation, especially in unstructured, cluttered, and semantically diverse environments. Recent researches have increasingly explored languag…

Learning Generalizable Feature Fields for Mobile Manipulation

2024-03-12 · Ri-Zhao Qiu, Yafei Hu, Yuchen Song, Ge Yang 외

An open problem in mobile manipulation is how to represent objects and scenes in a unified manner so that robots can use both for navigation and manipulation. The latter requires capturing intricate geometry while unders…

Novel View Synthesis