paper-with-me

홈 › Papers

VENOM: Versatile Embodied Network for Omni-bodied Motion tracking

2026-06-15 · Siddharth Padmanabhan, Kazuki Miyazawa, Takato Horii arxiv

Achieving expert-level expressive full-body motion tracking across multiple humanoids solely from demonstration data remains a challenging and relatively an underexplored problem in humanoid robot learning. Cross-embodiment motion tracking policies are mostly trained by decoupling the control problem into upper and lower body control. This work proposes VENOM, a cross-embodiment full-body motion tracking model for humanoids in simulation. VENOM is a GPT-based motion tracker trained on multiple humanoid data that can track the entire body without the requirement to split into upper and lower body control. We curate a multi-humanoid motion tracking dataset called the VENOM dataset that contains states, actions, and rewards and train VENOM and the baselines on this dataset. In this letter, we evaluate VENOM's performance against baselines and show that we can achieve a stable motion tracker across different humanoids more capable than an MLP trained on multiple humanoid data with supervised learning alone, and also show that despite lack of reward feedback, VENOM closely matches the tracking capability of experts that were trained using asymmetric-actor critic reinforcement learning.

📄 PDF Abstract BibTeX arXiv:2606.16696

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

OmniEVA: Embodied Versatile Planner via Task-Adaptive 3D-Grounded and Embodiment-aware Reasoning

2025-09-11 · Yuecheng Liu, Dafeng Chi, Shiguang Wu, Zhanguang Zhang 외 arxiv

Recent advances in multimodal large language models (MLLMs) have opened new opportunities for embodied intelligence, enabling multimodal understanding, reasoning, and interaction, as well as continuous spatial decision-m…

OmniActor: A Generalist GUI and Embodied Agent for 2D&3D Worlds

2025-09-02 · Longrong Yang, Zhixiong Zeng, Yufeng Zhong, Jing Huang 외 arxiv

Multimodal large language models are evolving toward multimodal agents capable of proactively executing tasks. Most agent research focuses on GUI or embodied scenarios, which correspond to agents interacting with 2D virt…

PANORAMA: The Rise of Omnidirectional Vision in the Embodied AI Era

2025-09-16 · Xu Zheng, Chenfei Liao, Ziqiao Weng, Kaiyu Lei 외 arxiv

Omnidirectional vision, using 360-degree vision to understand the environment, has become increasingly critical across domains like robotics, industrial inspection, and environmental monitoring. Compared to traditional p…

VLNVerse: A Benchmark for Vision-Language Navigation with Versatile, Embodied, Realistic Simulation and Evaluation

2025-12-22 · Sihao Lin, Zerui Li, Xunyi Zhao, Gengze Zhou 외 arxiv

Despite remarkable progress in Vision-Language Navigation (VLN), existing benchmarks remain confined to fixed, small-scale datasets with naive physical simulation. These shortcomings limit the insight that the benchmarks…

Vision-Language Navigation

OmniEcho: Spatial Audio Understanding for Embodied Agents

2026-09-20 · Ruixun Liu, Yuxuan Wang, Jiacheng Xie, Yuhuan You 외 hf

Humans can effortlessly localize the direction of a sound source and integrate it with visual cues for reasoning, yet this remains challenging for embodied agents. In particular, it is still unclear how to effectively ev…

Vision-Language Navigation