paper-with-me

홈 › Papers

Visually-grounded Humanoid Agents

2026-04-09 · Hang Ye, Xiaoxuan Ma, Fan Lu, Wayne Wu, Kwan-Yee Lin, Yizhou Wang arxiv

Digital human generation has been studied for decades and supports a wide range of real-world applications. However, most existing systems are passively animated, relying on privileged state or scripted control, which limits scalability to novel environments. We instead ask: how can digital humans actively behave using only visual observations and specified goals in novel scenes? Achieving this would enable populating any 3D environments with digital humans at scale that exhibit spontaneous, natural, goal-directed behaviors. To this end, we introduce Visually-grounded Humanoid Agents, a coupled two-layer (world-agent) paradigm that replicates humans at multiple levels: they look, perceive, reason, and behave like real people in real-world 3D scenes. The World Layer reconstructs semantically rich 3D Gaussian scenes from real-world videos via an occlusion-aware pipeline and accommodates animatable Gaussian-based human avatars. The Agent Layer transforms these avatars into autonomous humanoid agents, equipping them with first-person RGB-D perception and enabling them to perform accurate, embodied planning with spatial awareness and iterative reasoning, which is then executed at the low level as full-body actions to drive their behaviors in the scene. We further introduce a benchmark to evaluate humanoid-scene interaction in diverse reconstructed environments. Experiments show our agents achieve robust autonomous behavior, yielding higher task success rates and fewer collisions than ablations and state-of-the-art planning methods. This work enables active digital human population and advances human-centric embodied AI. Data, code, and models will be open-sourced.

📄 PDF Abstract BibTeX arXiv:2604.08509

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Embracing Evolution: A Call for Body-Control Co-Design in Embodied Humanoid Robot

2025-10-03 · Guiliang Liu, Bo Yue, Yi Jin Kim, Kui Jia arxiv

Humanoid robots, as general-purpose physical agents, must integrate both intelligent control and adaptive morphology to operate effectively in diverse real-world environments. While recent research has focused primarily …

HumanoidVLN: A Physics-Grounded Simulator and Benchmark for Vision-Language Navigation Across Diverse Humanoid Embodiments

2026-08-13 · Quan-Dung Pham, Anh Dao, The-Anh Nguyen, Minh Nguyen-Dinh 외 arxiv

Vision-Language Navigation (VLN) for humanoid robots poses challenges existing benchmarks fail to address: bipedal locomotion imposes physical constraints absent from wheeled agents, humanoid morphologies vary across pla…

Vision-Language NavigationReinforcement Learning

UniAct: Unified Motion Generation and Action Streaming for Humanoid Robots

2025-12-30 · Nan Jiang, Zimo He, Wanhe Yu, Lexi Pang 외 arxiv

A long-standing objective in humanoid robotics is the realization of versatile agents capable of following diverse multimodal instructions with human-level flexibility. Despite advances in humanoid control, bridging high…

Dialog without Dialog Data: Learning Visual Dialog Agents from VQA Data

2020-07-24 · NeurIPS 2020 12 · Michael Cogswell, Jiasen Lu, Rishabh Jain, Stefan Lee 외

Can we develop visually grounded dialog agents that can efficiently adapt to new tasks without forgetting how to talk to people? Such agents could leverage a larger variety of existing data to generalize to new tasks, mi…

Visual DialogVisual Question Answering (VQA)

A History-Aware Visually Grounded Critic for Computer Use Agents

2026-06-09 · Jaewoo Lee, Zaid Khan, Archiki Prasad, Justin Chih-Yao Chen 외 arxiv

Various test-time interventions for Computer Use Agents (CUAs), including critic models, have been developed to improve performance through pre-execution action evaluation in complex Graphical User Interface (GUI) enviro…

Visual Grounding