paper-with-me

Papers

Play2Perfect: What Matters in Dexterous Play Pretraining for Precise Assembly?

2026-06-24 · Tyler Ga Wei Lum, Kushal Kedia, C. Karen Liu, Jeannette Bohg hf

Multi-fingered robots promise the speed and dexterity of human hands, yet challenging problems such as precise assembly have remained out of reach. These tasks are contact-rich, making data collection for imitation learning difficult, and sparse-reward, making direct exploration with reinforcement learning (RL) intractable. Consequently, prior work has made progress by structuring the problem with specialized grippers, tool attachments, and environment fixtures. In this work, we argue that before a robot can perfect precise assembly, it must first learn to play. We further ask the question: what factors in the process of learning to play matter for precise assembly? We propose Play2Perfect, an RL framework for task-agnostic pretraining through play on diverse objects and goals, which is then perfected on precise assembly. The goal of play is to acquire reusable manipulation priors, such as grasping, in-hand reorientation and pose reaching. Finetuning then adapts this general prior to assembly, focusing exploration on the final contact-rich, high-precision interactions needed for success. We systematically study key design choices in play pretraining, including object diversity, training objective, trajectory diversity, and goal precision. We show that our prior is 33x more sample-efficient than RL training from scratch, even when provided with dense, multi-stage rewards. We demonstrate zero-shot sim-to-real transfer, achieving 60% success on tight insertions with only 0.5 mm contact clearance, and over 50% success on long-horizon multi-part assembly and screwing.

📄 PDF Abstract BibTeX arXiv:2606.26428

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Superhuman AI for Generals.io Using Self-Play Reinforcement Learning

2026-06-22 · Matej Straka, Viliam Lisý, Martin Schmid arxiv

We present a superhuman AI agent for Generals.io, a real-time strategy game that requires both long-horizon planning and short-term tactics under strong imperfect information. Trained for four days on 4x NVIDIA H200 GPUs…

Reinforcement Learning

Compromise, Don't Optimize: Generalizing Perfect Bayesian Equilibrium to Allow for Ambiguity

2020-03-05 · Karl Schlag, Andriy Zapechelnyuk

We introduce a solution concept for extensive-form games of incomplete information in which players need not assign likelihoods to what they do not know about the game. This is embedded in a model in which players can ho…

Enforcing Human-like Kinematics in Dexterous Piano Playing via Adversarial Posture Regularization

2026-06-22 · Bin Qiu, Yanming Shao, Guanyu Cai, Yao Mu arxiv

Reinforcement learning can train bimanual dexterous hands to play piano in physics simulation with high note accuracy, but for high-DoF dexterous hands, relying solely on task rewards or IK inversion often leads to unnat…

Reinforcement Learning

What Suppresses Nash Equilibrium Play in Large Language Models? Mechanistic Evidence and Causal Control

2026-04-29 · Paraskevas V. Lekeas, Giorgos Stamatopoulos arxiv

LLM agents are known to deviate from Nash equilibria in strategic interactions, but nobody has looked inside the model to understand why, or asked whether the deviation can be reversed. We do both. Working with four open…

Learning to Play Piano in the Real World

2025-03-19 · Yves-Simon Zeulner, Sandeep Selvaraj, Roberto Calandra

Towards the grand challenge of achieving human-level manipulation in robots, playing piano is a compelling testbed that requires strategic, precise, and flowing movements. Over the years, several works demonstrated hand-…