paper-with-me

Papers

SUGAR: A Scalable Human-Video-Driven Generalizable Humanoid Loco-Manipulation Learning Framework

2026-05-19 · Tianshu Wu, Xiangqi Kong, Yue Chen, Qize Yu, Hang Ye, Jia Li, Yizhou Wang, Hao Dong arxiv

Building humanoid robots capable of generalizable whole-body loco-manipulation in the real world remains a fundamental challenge. Existing methods either rely on laborious task-specific reward engineering, rigidly replay reference motions that fail to generalize, or depend on costly teleoperation that limits scalability. While human videos capture diverse human behaviors, motion priors inferred from them are inherently imperfect, suffering from occlusion, contact artifacts, and retargeting errors that render them unsuitable for direct policy learning. To address this, we present SUGAR, a scalable data-driven framework that converts diverse human videos into deployable humanoid loco-manipulation skills, without any task-specific reward engineering or reference-motion conditioning at inference. SUGAR proceeds in three stages. First, a fully automated pipeline extracts kinematic interaction priors including human-object motion trajectories and contact labels from unstructured human videos. Second, a privileged physics-based refiner uses a unified mimic reward and progressive state pool to transform imperfect priors into physically feasible, high-fidelity skills. Third, refined skills are distilled into a hierarchical autonomous policy consisting of a command generator and a command tracker. We evaluate SUGAR on six representative loco-manipulation tasks in simulation and real-world humanoid hardware. Our method substantially outperforms reference-tracking baselines, and performance scales clearly with the amount of human video data. It also achieves zero-shot real-world transfer with reliable closed-loop execution, autonomous failure recovery, and stable long-horizon performance under external perturbations. Project Page: https://tianshuwu.github.io/sugar-humanoid/

📄 PDF Abstract BibTeX arXiv:2605.20373

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner

2024-12-13 · Yufan Zhou, Ruiyi Zhang, Jiuxiang Gu, Nanxuan Zhao 외

We present SUGAR, a zero-shot method for subject-driven video customization. Given an input image, SUGAR is capable of generating videos for the subject contained in the image and aligning the generation with arbitrary v…

SUGAR: Pre-training 3D Visual Representations for Robotics

2024-04-01 · CVPR 2024 1 · ShiZhe Chen, Ricardo Garcia, Ivan Laptev, Cordelia Schmid

Learning generalizable visual representations from Internet data has yielded promising results for robotics. Yet, prevailing approaches focus on pre-training 2D representations, being sub-optimal to deal with occlusions …

3D Instance Segmentation3D Object RecognitionInstance SegmentationKnowledge Distillation+5

GIGA: Generalizable Sparse Image-driven Gaussian Avatars

2025-04-08 · Anton Zubekhin, Heming Zhu, Paulo Gotardo, Thabo Beeler 외

Driving a high-quality and photorealistic full-body human avatar, from only a few RGB cameras, is a challenging problem that has become increasingly relevant with emerging virtual reality technologies. To democratize suc…

SUGAR: A Sweeter Spot for Generative Unlearning of Many Identities

2025-12-06 · Dung Thuy Nguyen, Quang Nguyen, Preston K. Robinette, Eli Jiang 외 arxiv

Recent advances in 3D-aware generative models have enabled high-fidelity image synthesis of human identities. However, this progress raises urgent questions around user consent and the ability to remove specific individu…

HumanX: Toward Agile and Generalizable Humanoid Interaction Skills from Human Videos

2026-02-02 · Yinhuai Wang, Qihan Zhao, Yuen Fui Lau, Runyi Yu 외 arxiv

Enabling humanoid robots to perform agile and adaptive interactive tasks has long been a core challenge in robotics. Current approaches are bottlenecked by either the scarcity of realistic interaction data or the need fo…

Data Augmentation