paper-with-me

홈 › Papers

PhysDrift: Bridging the Embodiment Gap in Humanoid Co-Speech Motion Generation

2026-06-18 · Zhangzhao Liang, Xiaofen Xing, Mingyue Yang, Wenlve Zhou, Xiangmin Xu arxiv

Humanoid robots require co-speech motions that are not only expressive and speech-aligned, but also physically executable under embodiment constraints. Existing co-speech generation pipelines are predominantly human-centric: motions are first generated in human-body representations such as SMPL-X and subsequently retargeted to humanoid robots. In this work, we identify a fundamental embodiment gap in this paradigm, where the mismatch between human motion manifolds and humanoid embodiment constraints disrupts embodiment consistency during motion transfer and physical execution. Through extensive analysis, we show that although retargeting can preserve coarse motion semantics, it significantly compresses motion diversity and weakens prosody-motion synchronization, limiting expressive humanoid behaviors. To address this problem, we first propose IK-EER, a prosody-preserving humanoid motion curation framework that jointly optimizes kinematic feasibility and speech-motion temporal alignment during retargeting. Building upon the curated robot-native motion dataset, we further introduce PhysDrift, an embodiment-aware co-speech motion generation framework that directly predicts executable humanoid joint trajectories from speech without relying on intermediate human-body representations. Unlike conventional human-centric pipelines, PhysDrift maintains embodiment consistency throughout both training and inference while incorporating physical regularization to stabilize robot motion dynamics. Extensive experiments and real-world humanoid deployment demonstrate that embodiment-aware robot-native generation substantially improves speech-motion alignment, physical plausibility, motion smoothness, inference efficiency, and real-time interaction capability.

📄 PDF Abstract BibTeX arXiv:2606.19935

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Handroid: Bridging Dexterous Hand and Humanoid

2026-07-17 · Ruogu Li, Chenyang Ma, Sikai Li, Zhenyu Wei 외 arxiv

Dexterous hands and humanoid robots are typically developed as distinct embodiments: the former enable contact-rich manipulation at the object scale, whereas the latter provide mobility and whole-body interaction in huma…

OmniHumanoid: Streaming Cross-Embodiment Video Generation with Paired-Free Adaptation

2026-05-12 · Yiren Song, Xiyao Deng, Pei Yang, Yihan Wang 외 arxiv

Cross-embodiment video generation aims to transfer motions across different humanoid embodiments, such as human-to-robot and robot-to-robot, enabling scalable data generation for embodied intelligence. A major challenge …

Video Generation

Make Tracking Easy: Neural Motion Retargeting for Humanoid Whole-body Control

2026-03-23 · Qingrui Zhao, Kaiyue Yang, Xiyu Wang, Shiqi Zhao 외 arxiv

Humanoid robots require diverse motor skills to integrate into complex environments, but bridging the kinematic and dynamic embodiment gap from human data remains a major bottleneck. We demonstrate through Hessian analys…

Reinforcement Learning

H-Zero: Cross-Humanoid Locomotion Pretraining Enables Few-shot Novel Embodiment Transfer

2025-11-30 · Yunfeng Lin, Minghuan Liu, Yufei Xue, Ming Zhou 외 arxiv

The rapid advancement of humanoid robotics has intensified the need for robust and adaptable controllers to enable stable and efficient locomotion across diverse platforms. However, developing such controllers remains a …

Semantic Co-Speech Gesture Synthesis and Real-Time Control for Humanoid Robots

2025-12-19 · Gang Zhang arxiv

We present an innovative end-to-end framework for synthesizing semantically meaningful co-speech gestures and deploying them in real-time on a humanoid robot. This system addresses the challenge of creating natural, expr…

Gesture Generation