paper-with-me

홈 › Papers

EmbodMocap: In-the-Wild 4D Human-Scene Reconstruction for Embodied Agents

2026-02-26 · Wenjia Wang, Liang Pan, Huaijin Pi, Yuke Lou, Xuqian Ren, Yifan Wu, Zhouyingcheng Liao, Lei Yang, Rishabh Dabral, Christian Theobalt, Taku Komura arxiv

Human behaviors in the real world naturally encode rich, long-term contextual information that can be leveraged to train embodied agents for perception, understanding, and acting. However, existing capture systems typically rely on costly studio setups and wearable devices, limiting the large-scale collection of scene-conditioned human motion data in the wild. To address this, we propose EmbodMocap, a portable and affordable data collection pipeline using two moving iPhones. Our key idea is to jointly calibrate dual RGB-D sequences to reconstruct both humans and scenes within a unified metric world coordinate frame. The proposed method allows metric-scale and scene-consistent capture in everyday environments without static cameras or markers, bridging human motion and scene geometry seamlessly. Compared with optical capture ground truth, we demonstrate that the dual-view setting exhibits a remarkable ability to mitigate depth ambiguity, achieving superior alignment and reconstruction performance over single iphone or monocular models. Based on the collected data, we empower three embodied AI tasks: monocular human-scene-reconstruction, where we fine-tune on feedforward models that output metric-scale, world-space aligned humans and scenes; physics-based character animation, where we prove our data could be used to scale human-object interaction skills and scene-aware motion tracking; and robot motion control, where we train a humanoid robot via sim-to-real RL to replicate human motions depicted in videos. Experimental results validate the effectiveness of our pipeline and its contributions towards advancing embodied AI research.

📄 PDF Abstract BibTeX arXiv:2602.23205

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

WildCity: A Real-World City-Scale Testbed for Rendering, Simulation, and Spatial Intelligence

2026-07-07 · Xiangyu Han, Mengyu Yang, Jiaqi Li, Bowen Chang 외 arxiv

Humans can navigate an unfamiliar city and gradually form a coherent spatial mental map spanning tens of square kilometers. Can AI build spatial representations at a comparable scale? Although recent foundation models ha…

Hand3R: Online 4D Hand-Scene Reconstruction in the Wild

2026-02-03 · Wendi Hu, Haonan Zhou, Wenhao Hu, Gaoang Wang arxiv

For Embodied AI, jointly reconstructing dynamic hands and the dense scene context is crucial for understanding physical interaction. However, most existing methods recover isolated hands in local coordinates, overlooking…

Total-Recon: Deformable Scene Reconstruction for Embodied View Synthesis

2023-04-24 · ICCV 2023 1 · Chonghyuk Song, Gengshan Yang, Kangle Deng, Jun-Yan Zhu 외

We explore the task of embodied view synthesis from monocular videos of deformable scenes. Given a minute-long RGBD video of people interacting with their pets, we render the scene from novel camera trajectories derived …

Joint Optimization for 4D Human-Scene Reconstruction in the Wild

2025-01-04 · Zhizheng Liu, Joe Lin, Wayne Wu, Bolei Zhou

Reconstructing human motion and its surrounding environment is crucial for understanding human-scene interaction and predicting human movements in the scene. While much progress has been made in capturing human-scene int…

Human Mesh RecoveryMotion Estimation

EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI

2023-12-26 · CVPR 2024 1 · Tai Wang, Xiaohan Mao, Chenming Zhu, Runsen Xu 외

In the realm of computer vision and robotics, embodied agents are expected to explore their environment and carry out human instructions. This necessitates the ability to fully understand 3D scenes given their first-pers…

Scene Understanding