paper-with-me

Papers

Mitty: Diffusion-based Human-to-Robot Video Generation

2025-12-19 · Yiren Song, Cheng Liu, Weijia Mao, Mike Zheng Shou arxiv

Learning directly from human demonstration videos is a key milestone toward scalable and generalizable robot learning. Yet existing methods rely on intermediate representations such as keypoints or trajectories, introducing information loss and cumulative errors that harm temporal and visual consistency. We present Mitty, a Diffusion Transformer that enables video In-Context Learning for end-to-end Human2Robot video generation. Built on a pretrained video diffusion model, Mitty leverages strong visual-temporal priors to translate human demonstrations into robot-execution videos without action labels or intermediate abstractions. Demonstration videos are compressed into condition tokens and fused with robot denoising tokens through bidirectional attention during diffusion. To mitigate paired-data scarcity, we also develop an automatic synthesis pipeline that produces high-quality human-robot pairs from large egocentric datasets. Experiments on Human2Robot and EPIC-Kitchens show that Mitty delivers state-of-the-art results, strong generalization to unseen environments, and new insights for scalable robot learning from human observations.

📄 PDF Abstract BibTeX arXiv:2512.17253

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

VimoRAG: Video-based Retrieval-augmented 3D Motion Generation for Motion Language Models

2025-08-16 · Haidong Xu, Guangwei Xu, Zhedong Zheng, Xiatian Zhu 외 arxiv

This paper introduces VimoRAG, a novel video-based retrieval-augmented motion generation framework for motion large language models (LLMs). As motion LLMs face severe out-of-domain/out-of-vocabulary issues due to limited…

Video Retrieval

Action Agent: Agentic Video Generation Meets Flow-Constrained Diffusion

2026-05-02 · Jeffrin Sam, Nguyen Khang, Yara Mahmoud, Miguel Altamirano Cabrera 외 arxiv

We present Action Agent, a two-stage framework that unifies agentic navigation video generation with flow-constrained diffusion control for multi-embodiment robot navigation. In Stage I, a large language model (LLM) acts…

Video GenerationRobot Navigation

Learning an Actionable Discrete Diffusion Policy via Large-Scale Actionless Video Pre-Training

2024-02-22 · Haoran He, Chenjia Bai, Ling Pan, Weinan Zhang 외

Learning a generalist embodied agent capable of completing multiple tasks poses challenges, primarily stemming from the scarcity of action-labeled robotic datasets. In contrast, a vast amount of human videos exist, captu…

CRAFT: Video Diffusion for Bimanual Robot Data Generation

2026-04-04 · Jason Chen, I-Chun Arthur Liu, Gaurav Sukhatme, Daniel Seita arxiv

Bimanual robot learning from demonstrations is fundamentally limited by the cost and narrow visual diversity of real-world data, which constrains policy robustness across viewpoints, object configurations, and embodiment…

Video Generation

Virtual avatar generation models as world navigators

2024-06-03 · Sai Mandava

We introduce SABR-CLIMB, a novel video model simulating human movement in rock climbing environments using a virtual avatar. Our diffusion transformer predicts the sample instead of noise in each diffusion step and inges…