paper-with-me

홈 › Papers

HACMan++: Spatially-Grounded Motion Primitives for Manipulation

2024-07-11 · Bowen Jiang, Yilin Wu, Wenxuan Zhou, Chris Paxton, David Held

Although end-to-end robot learning has shown some success for robot manipulation, the learned policies are often not sufficiently robust to variations in object pose or geometry. To improve the policy generalization, we introduce spatially-grounded parameterized motion primitives in our method HACMan++. Specifically, we propose an action representation consisting of three components: what primitive type (such as grasp or push) to execute, where the primitive will be grounded (e.g. where the gripper will make contact with the world), and how the primitive motion is executed, such as parameters specifying the push direction or grasp orientation. These three components define a novel discrete-continuous action space for reinforcement learning. Our framework enables robot agents to learn to chain diverse motion primitives together and select appropriate primitive parameters to complete long-horizon manipulation tasks. By grounding the primitives on a spatial location in the environment, our method is able to effectively generalize across object shape and pose variations. Our approach significantly outperforms existing methods, particularly in complex scenarios demanding both high-level sequential reasoning and object generalization. With zero-shot sim-to-real transfer, our policy succeeds in challenging real-world manipulation tasks, with generalization to unseen objects. Videos can be found on the project website: https://sgmp-rss2024.github.io.

📄 PDF Abstract BibTeX arXiv:2407.08585

Code (0)

등록된 구현이 없습니다.

Tasks

ObjectRobot Manipulation

Similar Papers 제목 키워드 기반

HACMan: Learning Hybrid Actor-Critic Maps for 6D Non-Prehensile Manipulation

2023-05-06 · Wenxuan Zhou, Bowen Jiang, Fan Yang, Chris Paxton 외

Manipulating objects without grasping them is an essential component of human dexterity, referred to as non-prehensile manipulation. Non-prehensile manipulation may enable more complex interactions with the objects, but …

Object

Spatially Grounded Long-Horizon Task Planning in the Wild

2026-03-13 · Sehun Jung, HyunJee Song, Dong-Hee Kim, Reuben Tan 외 arxiv

Recent advances in robot manipulation increasingly leverage Vision-Language Models (VLMs) for high-level reasoning, such as decomposing task instructions into sequential action plans expressed in natural language that gu…

Robot Manipulation

SG-VLA: Learning Spatially-Grounded Vision-Language-Action Models for Mobile Manipulation

2026-03-24 · Ruisen Tu, Arth Shukla, Sohyun Yoo, Xuanlin Li 외 arxiv

Vision-Language-Action (VLA) models show promise for robotic control, yet performance in complex household environments remains sub-optimal. Mobile manipulation requires reasoning about global scene layout, fine-grained …

Language-Grounded Decoupled Action Representation for Robotic Manipulation

2026-03-13 · Wuding Weng, Tongshu Wu, Liucheng Chen, Siyu Xie 외 arxiv

The heterogeneity between high-level vision-language understanding and low-level action control remains a fundamental challenge in robotic manipulation. Although recent methods have advanced task-specific action alignmen…

Contrastive Learning

Ground4D: Spatially-Grounded Feedforward 4D Reconstruction for Unstructured Off-Road Scenes

2026-05-06 · Shuo Wang, Jilin Mei, Fuyang Liu, Wenfei Guan 외 arxiv

Feedforward Gaussian Splatting has recently emerged as an efficient paradigm for 4D reconstruction in autonomous driving. However, in unstructured off-road scenes, its performance degrades due to high-frequency geometry,…

Autonomous Driving