paper-with-me

Papers

JUICER: Data-Efficient Imitation Learning for Robotic Assembly

2024-04-04 · Lars Ankile, Anthony Simeonov, Idan Shenfeld, Pulkit Agrawal

While learning from demonstrations is powerful for acquiring visuomotor policies, high-performance imitation without large demonstration datasets remains challenging for tasks requiring precise, long-horizon manipulation. This paper proposes a pipeline for improving imitation learning performance with a small human demonstration budget. We apply our approach to assembly tasks that require precisely grasping, reorienting, and inserting multiple parts over long horizons and multiple task phases. Our pipeline combines expressive policy architectures and various techniques for dataset expansion and simulation-based data augmentation. These help expand dataset support and supervise the model with locally corrective actions near bottleneck regions requiring high precision. We demonstrate our pipeline on four furniture assembly tasks in simulation, enabling a manipulator to assemble up to five parts over nearly 2500 time steps directly from RGB images, outperforming imitation and data augmentation baselines. Project website: https://imitation-juicer.github.io/.

📄 PDF Abstract BibTeX arXiv:2404.03729

Code (1)

ankile/imitation-juicer 공식 구현 pytorch

Tasks

Data AugmentationImitation Learning

Similar Papers 제목 키워드 기반

VLM-driven Skill Selection for Robotic Assembly Tasks

2025-11-07 · Jeong-Jung Kim, Doo-Yeol Koh, Chang-Hyun Kim arxiv

This paper presents a robotic assembly framework that combines Vision-Language Models (VLMs) with imitation learning for assembly manipulation tasks. Our system employs a gripper-equipped robot that moves in 3D space to …

Natural Language Understanding

Data-Juicer: A One-Stop Data Processing System for Large Language Models

2023-09-05 · Daoyuan Chen, Yilun Huang, Zhijian Ma, Hesen Chen 외

The immense evolution in Large Language Models (LLMs) has underscored the importance of massive, heterogeneous, and high-quality data. A data recipe is a mixture of data from different sources for training LLMs, which pl…

Distributed Computing

Vision-Language-Action Models for Selective Robotic Disassembly: A Case Study on Critical Component Extraction from Desktops

2025-12-04 · Chang Liu, Sibo Tian, Sara Behdad, Xiao Liang 외 arxiv

Automating disassembly of critical components from end-of-life (EoL) desktops, such as high-value items like RAM modules and CPUs, as well as sensitive parts like hard disk drives, remains challenging due to the inherent…

Motion Planning

AssemLM: A Spatial Reasoning Multimodal Large Language Model for Robotic Assembly

2026-04-10 · Zhi Jing, Jinbin Qiao, Ouyang Lu, Jicong Ao 외 arxiv

Spatial reasoning is a fundamental capability for embodied intelligence, especially for fine-grained manipulation tasks such as robotic assembly. Recent methods based on vision-language models (VLMs) largely rely on coar…

Spatial ReasoningPoint Clouds

WorkBenchMark: A LEGO-Based Assembly Benchmark with an Assembly-by-Disassembly Baseline for the Smart Manufacturing League

2026-06-02 · Wenbo Ma, Daniel Swoboda, Matteo Tschesche, Till Hofmann arxiv

We introduceWorkBenchMark, a LEGO Duplo-based robotic assembly benchmark motivated by the RoboCup Smart Manufacturing League. Robotic assembly couples low-level manipulation with task-level symbolic reasoning under physi…