paper-with-me

Papers

Closing the Loop in Humanoid VLA: Persistent 3D Object Tokens for Verifiable Loco-Manipulation

2026-07-20 · Peng Ren, Haoyang Ge, Jiang Zhao, Cong Huang, Yukun Shi, Pei Chi, Kai Chen arxiv

Vision-language-action policies are a promising foundation for general robot control, but long-horizon humanoid loco-manipulation requires the robot to treat task objects as persistent physical entities across movement, contact, occlusion, and recovery. We study this problem as object-state divergence: the object state used to condition a whole-body action can differ from the state used to decide whether the action achieved the intended physical relation. We propose \emph{Persistent Object Tokenization} (POT), which maintains role-indexed 3D object records from RGB-D observations and converts them into object tokens for a whole-body action expert. Instantiated as \emph{POT-VLA}, the same object records condition action generation and support geometric predicate checks, yielding a closed-loop execution system in which object state is both actionable and verifiable. On a Unitree G1, POT-VLA improves a matched direct GR00T-N1.7 baseline from 39/80 to 71/80 successes over eight real-world task families. In an external Being-0-aligned reference, POT-VLA achieves 44/50 successes on aligned service tasks, compared with the 37/50 success reported by the Being-0 paper. The largest gains occur on tasks requiring maintained 3D relations, suggesting that persistent object-centered state is a useful abstraction for verifiable humanoid VLA execution.

📄 PDF Abstract BibTeX arXiv:2607.18016

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Pro-HOI: Perceptive Root-guided Humanoid-Object Interaction

2026-03-01 · Yuhang Lin, Jiyuan Shi, Dewei Wang, Jipeng Kong 외 arxiv

Executing reliable Humanoid-Object Interaction (HOI) tasks for humanoid robots is hindered by the lack of generalized control interfaces and robust closed-loop perception mechanisms. In this work, we introduce Perceptive…

Emergent Active Perception and Dexterity of Simulated Humanoids from Visual Reinforcement Learning

2025-05-18 · Zhengyi Luo, Chen Tessler, Toru Lin, Ye Yuan 외

Human behavior is fundamentally shaped by visual perception -- our ability to interact with the world depends on actively gathering relevant information and adapting our movements accordingly. Behaviors like searching fo…

Object

Learning to Evolve: Multi-modal Interactive Fields for Robust Humanoid Navigation in Dynamic Environments

2026-05-21 · Peifeng Jiang, Hong Liu, Jin Jin, Wenshuai Wang 외 arxiv

Safe manipulation-oriented navigation for humanoid robots requires scene memory that remains reliable under locomotion-induced perceptual distortion, environmental changes, and interaction-level geometric safety constrai…

VOFA: Visual Object Goal Pushing with Force-Adaptive Control for Humanoids

2026-05-02 · Zichao Hu, Zifan Xu, Dongsik Chang, He Yin 외 arxiv

The ability to push large objects in a goal-directed manner using onboard egocentric perception is an essential skill for humanoid robots to perform complex tasks such as material handling in warehouses. To robustly mani…

RoboReact: Agentic Skill Distillation from Generated Egocentric Videos for Generalizable Whole-Body Manipulation

2026-08-04 · Shuliang He, Shuai Wang, Bo Yue, Junchi Teng 외 arxiv

Humanoid robots have the potential to perform dexterous manipulation in human environments, yet acquiring diverse and generalizable skills remains costly due to expensive hardware data collection and labor-intensive anno…

3D Reconstruction