paper-with-me

홈 › Papers

Transferring the Intelligence of VLMs to Robotic Control

2026-09-19 · Meng-Hao Guo, Zhe-Han Mo, Jia-Jun Wang, Yi Zhang, Kejin Wang, Yi-Xuan Deng, Jia-Peng Zhang, Yongming Rao, Shi-Min Hu hf

Humans can seamlessly adapt to both physical and digital worlds, suggesting that while a digital-to-real gap exists in embodiment, environment and task, human intelligence itself may transfer across this gap. This naturally raises a fundamental question: can the intelligence of vision-language models (VLMs) similarly generalize from the digital world to the physical world for robotic control? We investigate this question through RoboDawn, a human-intuitive interface that exposes robotic control to an agentic VLM through a compact set of discrete translation, rotation, and gripper commands. Using this interface, the VLM controls a robot in a closed loop: it observes the current visual state, reasons about the next action, executes it, and adapts subsequent decisions to the resulting state. Furthermore, we introduce an in-context learning (ICL) scheme that uses a few demonstrations to ground the VLM in both interface usage and task-solving strategies. Experiments on RoboTwin 2.0 C2R and RoboDojo demonstrate that RoboDawn achieves strong performance without task-specific robot training. In the zero-shot setting, RoboDawn outperforms several strong policies trained on benchmarkspecific robot data, while a single in-context demonstration further yields substantial performance gains and establishes state-of-the-art (SOTA) results. On RoboTwin 2.0 C2R, the success rate increases from 53.2% zero-shot to 73.6% one-shot, exceeding the solid baseline π0.5 (46.0%). Similar gains are observed on RoboDojo, where success rate improves from 35.67% zero-shot to 47.17% one-shot. The same framework also transfers to real-world robots, performing block-in-basket and block stacking on Franka.

📄 PDF Abstract BibTeX arXiv:2609.22966

Code (1)

Valiant-Cat/hfpaper

Similar Papers 제목 키워드 기반

Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes

2025-10-22 · Zhiyuan Feng, Zhaolu Kang, Qijie Wang, Zhiying Du 외 arxiv

Vision-language models (VLMs) are essential to Embodied AI, enabling robots to perceive, reason, and act in complex environments. They also serve as the foundation for the recent Vision-Language-Action (VLA) models. Yet …

Spatial Reasoning

TwinBrainVLA: Unleashing the Potential of Generalist VLMs for Embodied Tasks via Asymmetric Mixture-of-Transformers

2026-01-20 · Bin Yu, Shijie Lian, Xiaopeng Lin, Yuliang Wei 외 arxiv

The fundamental premise of Vision-Language-Action (VLA) models is to harness the extensive general capabilities of pre-trained Vision-Language Models (VLMs) for generalized embodied intelligence. However, standard roboti…

Continuous Control

ManipBench: Benchmarking Vision-Language Models for Low-Level Robot Manipulation

2025-05-14 · Enyu Zhao, Vedant Raval, Hejia Zhang, Jiageng Mao 외

Vision-Language Models (VLMs) have revolutionized artificial intelligence and robotics due to their commonsense reasoning capabilities. In robotic manipulation, VLMs are used primarily as high-level planners, but recent …

BenchmarkingDeformable Object ManipulationObjectRobot Manipulation

Abstract 3D Perception for Spatial Intelligence in Vision-Language Models

2025-11-14 · Yifan Liu, Fangneng Zhan, Kaichen Zhou, Yilun Du 외 arxiv

Vision-language models (VLMs) struggle with 3D-related tasks such as spatial cognition and physical understanding, which are crucial for real-world applications like robotics and embodied agents. We attribute this to a m…

PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

2024-02-12 · Soroush Nasiriany, Fei Xia, Wenhao Yu, Ted Xiao 외

Vision language models (VLMs) have shown impressive capabilities across a variety of tasks, from logical reasoning to visual understanding. This opens the door to richer interaction with the world, for example robotic co…

Instruction FollowingLogical ReasoningQuestion AnsweringSpatial Reasoning+2