paper-with-me

홈 › Papers

From Passive Video to Editable Experience: Physically Grounded Experience Synthesis for Embodied Intelligence

2026-07-29 · Jia Luo arxiv

The key bottleneck in embodied AI is not model architecture but data. Although billions of human manipulation videos exist online, robots cannot directly learn from them due to the embodiment gap between human morphology and robot hardware. We introduce Pegasus, a low-resource framework that bridges this gap by translating human demonstrations into robot-learnable data through structured knowledge transfer. Instead of relying on raw video prompts, Pegasus constructs a graph-based intermediate representation: a Task Graph extracted from human videos is transformed through Affordance and Constraint Graphs into a Robot Planning Graph for robot-conditioned video generation. A hierarchical affordance latent space models the relationship between object states, affordances, and tasks, enabling generalization beyond object identities. A closed-loop physics verifier further filters invalid generations using kinematic feasibility, collision constraints, and joint limits. We evaluate Pegasus across a range of egocentric manipulation benchmarks, including GTEA Gaze+ and EPIC-KITCHENS-100, and diverse robot embodiments, assessing Task Correctness, Executability, State Consistency, and Learnability. Results demonstrate reliable cross-embodiment translation and show that robot data generation can be reframed from a hardware collection problem into a scalable, low-resource knowledge transfer problem.

📄 PDF Abstract BibTeX arXiv:2607.26903

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

Mirage2Matter: A Physically Grounded Gaussian World Model from Video

2026-01-24 · Zhengqing Gao, Ziwen Li, Xin Wang, Jiaxin Huang 외 arxiv

The scalability of embodied intelligence is fundamentally constrained by the scarcity of real-world interaction data. While simulation platforms provide a promising alternative, existing approaches often suffer from a su…

V-RGBX: Video Editing with Accurate Controls over Intrinsic Properties

2025-12-12 · Ye Fang, Tong Wu, Valentin Deschaintre, Duygu Ceylan 외 arxiv

Large-scale video generation models have shown remarkable potential in modeling photorealistic appearance and lighting interactions in real-world scenes. However, a closed-loop framework that jointly understands intrinsi…

Inverse RenderingVideo Generation

ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs

2026-09-09 · Yizhan Li, Jianxin You, Mengyang Xiong, Yinhuan Chen 외 hf

Reacting to sudden physical hazards (catching a slipping plate, dodging a falling knife) is both a meaningful test of embodied intelligence and a hard requirement for deploying multimodal large language models (MLLMs) as…

Question Answering

MultiGen: Level-Design for Editable Multiplayer Worlds in Diffusion Game Engines

2026-03-03 · Ryan Po, David Junhao Zhang, Amir Hertz, Gordon Wetzstein 외 arxiv

Video world models have shown immense promise for interactive simulation and entertainment, but current systems still struggle with two important aspects of interactivity: user control over the environment for reproducib…

ObjectForesight: Predicting Future 3D Object Trajectories from Human Videos

2026-01-08 · Rustin Soraki, Homanga Bharadhwaj, Ali Farhadi, Roozbeh Mottaghi arxiv

Humans can effortlessly anticipate how objects might move or change through interaction--imagining a cup being lifted, a knife slicing, or a lid being closed. We aim to endow computational systems with a similar ability …

3D Pose Estimation