paper-with-me

Papers

HumanCLAW: Can Vision-Language Models Act Through a Body?

2026-07-29 · Siyao Li, Jiawei Gu, Shuai Liu, Kairui Hu, Zekun Li, Linjie Li, Chengcheng Tang, Po-Chen Wu, Ivan Shugurov, Lingni Ma, Michael Zollhoefer, Sizhe An, Abhay Mittal, Amy Zhao, Ranjay Krishna, Manling Li, Ziwei Liu, Chuan Guo hf

Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome of an action couples the VLM's decision with motor control. When a task fails, it is hard to tell whether the VLM made a bad choice or the motor controller simply failed to execute it, e.g., losing balance and falling. In this work, we introduce HumanCLAW, an evaluation framework that decouples action decision-making from low-level execution. At every step, a harnessed, off-the-shelf VLM issues an atomic skill command, and the command is translated into a sub-second chunk of continuous full-body motion with real physical consequences, including gravity and collisions. The body can therefore act freely in the physical world, while execution-side disturbances, balance and motor errors, are factored out. What remains measurable is the model's action intelligence: its moment-to-moment choice of what the body should execute next. Based on this framework, we build HumanCLAW-Bench: 1,218 long-horizon, egocentric find-navigate-interact episodes across 41 indoor scenes. We test nine state-of-the-art VLMs and find that none solves the benchmark; the best model reaches only a 16.8% success rate. Recognizing the target is not the bottleneck. What current VLMs lack is embodied self-awareness: they lose track of their own body, failing to tell where it is, whether it has reached the goal, or whether it has hit an obstacle.

📄 PDF Abstract BibTeX arXiv:2607.27180

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model

2026-09-08 · Anqi Li, Yuxin Chen, Zhaobo Li, Zhuo Cao 외 hf

We study the problem of navigating cluttered indoor environments with a humanoid robot. Unlike conventional methods that model navigation as a 2D path planning problem, humanoid traversal in cluttered environments requir…

Vision-Language Navigation

SWIM: Vision-Language-Grounded Soft Whole-Body Interactive Manipulation

2026-09-15 · Tingcong Liu, Aye Phyu Phyu Aung, Junjie Xiong, Siyi Ma 외 arxiv

Soft and continuum robots enable manipulation through distributed body deformation and contact, yet translating language and visual context into executable whole-body actuation remains a fundamental challenge. We present…

TWIST2: Scalable, Portable, and Holistic Humanoid Data Collection System

2025-11-04 · Yanjie Ze, Siheng Zhao, Weizhuo Wang, Angjoo Kanazawa 외 arxiv

Large-scale data has driven breakthroughs in robotics, from language models to vision-language-action models in bimanual manipulation. However, humanoid robotics lacks equally effective data collection frameworks. Existi…

Real-time Bangla Sign Language Translator

2024-12-21 · Rotan Hawlader Pranto, Shahnewaz Siddique

The human body communicates through various meaningful gestures, with sign language using hands being a prominent example. Bangla Sign Language Translation (BSLT) aims to bridge communication gaps for the deaf and mute c…

Sign Language TranslationTranslation

LeVERB: Humanoid Whole-Body Control with Latent Vision-Language Instruction

2025-06-16 · Haoru Xue, Xiaoyu Huang, Dantong Niu, Qiayuan Liao 외

Vision-language-action (VLA) models have demonstrated strong semantic understanding and zero-shot generalization, yet most existing systems assume an accurate low-level controller with hand-crafted action "vocabulary" su…

Instruction FollowingVision-Language-ActionVisual NavigationZero-shot Generalization