paper-with-me

홈 › Papers

Manual2Skill++: Connector-Aware General Robotic Assembly from Instruction Manuals via Vision-Language Models

2025-10-18 · Chenrui Tie, Shengxiang Sun, Yudi Lin, Yanbo Wang, Zhongrui Li, Zhouhan Zhong, Jinxuan Zhu, Yiman Pang, Haonan Chen, Junting Chen, Ruihai Wu, Lin Shao arxiv

Assembly hinges on reliably forming connections between parts; yet most robotic approaches plan assembly sequences and part poses while treating connectors as an afterthought. Connections represent the foundational physical constraints of assembly execution; while task planning sequences operations, the precise establishment of these constraints ultimately determines assembly success. In this paper, we treat connections as explicit, primary entities in assembly representation, directly encoding connector types, specifications, and locations for every assembly step. Drawing inspiration from how humans learn assembly tasks through step-by-step instruction manuals, we present Manual2Skill++, a vision-language framework that automatically extracts structured connection information from assembly manuals. We encode assembly tasks as hierarchical graphs where nodes represent parts and sub-assemblies, and edges explicitly model connection relationships between components. A large-scale vision-language model parses symbolic diagrams and annotations in manuals to instantiate these graphs, leveraging the rich connection knowledge embedded in human-designed instructions. We curate a dataset containing over 20 assembly tasks with diverse connector types to validate our representation extraction approach, and evaluate the complete task understanding-to-execution pipeline across four complex assembly scenarios in simulation, spanning furniture, toys, and manufacturing components with real-world correspondence. More detailed information can be found at https://nus-lins-lab.github.io/Manual2SkillPP/

📄 PDF Abstract BibTeX arXiv:2510.16344

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Uni-Skill: Building Self-Evolving Skill Repository for Generalizable Robotic Manipulation

2026-03-03 · Senwei Xie, Yuntian Zhang, Ruiping Wang, Xilin Chen arxiv

While skill-centric approaches leverage foundation models to enhance generalization in compositional tasks, they often rely on fixed skill libraries, limiting adaptability to new tasks without manual intervention. To add…

Zero-shot Generalization

Being-0: A Humanoid Robotic Agent with Vision-Language Models and Modular Skills

2025-03-16 · Haoqi Yuan, Yu Bai, Yuhui Fu, Bohan Zhou 외

Building autonomous robotic agents capable of achieving human-level performance in real-world embodied tasks is an ultimate goal in humanoid robot research. Recent advances have made significant progress in high-level co…

Task Planning

Skill-Aware Diffusion for Generalizable Robotic Manipulation

2026-01-16 · Aoshen Huang, Jiaming Chen, Jiyu Cheng, Ran Song 외 arxiv

Robust generalization in robotic manipulation is crucial for robots to adapt flexibly to diverse environments. Existing methods usually improve generalization by scaling data and networks, but model tasks independently a…

Touch2Insert: Zero-Shot Peg Insertion by Touching Intersections of Peg and Hole

2026-03-04 · Masaru Yajima, Yuma Shin, Rei Kawakami, Asako Kanezaki 외 arxiv

Reliable insertion of industrial connectors remains a central challenge in robotics, requiring sub-millimeter precision under uncertainty and often without full visual access. Vision-based approaches struggle with occlus…

Pose Estimation

MOSAIC: Skill-Centric Manipulation Planning with Physics Simulation

2025-04-23 · Itamar Mishani, Yorai Shaoul, Maxim Likhachev arxiv

Planning long-horizon manipulation motions using a set of predefined skills is a central challenge in robotics; solving it efficiently could enable general-purpose robots to tackle novel tasks by flexibly composing gener…

Motion Planning