paper-with-me

홈 › Papers

Tenma: Robust Cross-Embodiment Robot Manipulation with Diffusion Transformer

2025-09-15 · Travis Davies, Yiqi Huang, Yunxin Liu, Xiang Chen, Huxian Liu, Luhui Hu arxiv

Scaling Transformer policies and diffusion models has advanced robotic manipulation, yet combining these techniques in lightweight, cross-embodiment learning settings remains challenging. We study design choices that most affect stability and performance for diffusion-transformer policies trained on heterogeneous, multimodal robot data, and introduce Tenma, a lightweight diffusion-transformer for bi-manual arm control. Tenma integrates multiview RGB, proprioception, and language via a cross-embodiment normalizer that maps disparate state/action spaces into a shared latent space; a Joint State-Time encoder for temporally aligned observation learning with inference speed boosts; and a diffusion action decoder optimized for training stability and learning capacity. Across benchmarks and under matched compute, Tenma achieves an average success rate of 88.95% in-distribution and maintains strong performance under object and scene shifts, substantially exceeding baseline policies whose best in-distribution average is 18.12%. Despite using moderate data scale, Tenma delivers robust manipulation and generalization, indicating the great potential for multimodal and cross-embodiment learning strategies for further augmenting the capacity of transformer-based imitation learning policies.

📄 PDF Abstract BibTeX arXiv:2509.11865

Code (0)

등록된 구현이 없습니다.

Tasks

Robot Manipulation

Similar Papers 제목 키워드 기반

H-RDT: Human Manipulation Enhanced Bimanual Robotic Manipulation

2025-07-31 · Hongzhe Bi, Lingxuan Wu, Tianwei Lin, Hengkai Tan 외 arxiv

Imitation learning for robotic manipulation faces a fundamental challenge: the scarcity of large-scale, high-quality robot demonstration data. Recent robotic foundation models often pre-train on cross-embodiment robot da…

Robot ManipulationFew-Shot Learning

X-Diffusion: Training Diffusion Policies on Cross-Embodiment Human Demonstrations

2025-11-06 · Maximus A. Pace, Prithwish Dan, Chuanruo Ning, Atiksh Bhardwaj 외 arxiv

Human videos are a scalable source of training data for robot learning. However, humans and robots significantly differ in embodiment, making many human actions infeasible for direct execution on a robot. Still, these de…

UMI-on-Air: Embodiment-Aware Guidance for Embodiment-Agnostic Visuomotor Policies

2025-10-02 · Harsh Gupta, Xiaofeng Guo, Huy Ha, Chuer Pan 외 arxiv

We introduce UMI-on-Air, a framework for embodiment-aware deployment of embodiment-agnostic manipulation policies. Our approach leverages diverse, unconstrained human demonstrations collected with a handheld gripper (UMI…

Transformer Transformer: A Unified Model for Motion-Conditioned Robot Co-design

2026-07-28 · Huy Ha, C. Karen Liu, Shuran Song arxiv

An often overlooked factor of robot manipulation performance is the embodiment of the robot itself. Motivated by this problem, we study motion-conditioned robot co-design, where the goal is to generate complete robot des…

Robot Manipulation

Physics-Driven Data Generation for Contact-Rich Manipulation via Trajectory Optimization

2025-02-27 · Lujie Yang, H. J. Terry Suh, Tong Zhao, Bernhard Paus Graesdal 외

We present a low-cost data generation pipeline that integrates physics-based simulation, human demonstrations, and model-based planning to efficiently generate large-scale, high-quality datasets for contact-rich robotic …

Contact-rich Manipulation