paper-with-me

Papers

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

2025-05-30 · Gen Luo, Ganlin Yang, Ziyang Gong, Guanzhou Chen, Haonan Duan, Erfei Cui, Ronglei Tong, Zhi Hou, Tianyi Zhang, Zhe Chen, Shenglong Ye, Lewei Lu, Jingbo Wang, Wenhai Wang, Jifeng Dai, Yu Qiao, Rongrong Ji, Xizhou Zhu

The remarkable progress of Multimodal Large Language Models (MLLMs) has attracted increasing attention to extend them to physical entities like legged robot. This typically requires MLLMs to not only grasp multimodal understanding abilities, but also integrate visual-spatial reasoning and physical interaction capabilities. Nevertheless,existing methods struggle to unify these capabilities due to their fundamental differences.In this paper, we present the Visual Embodied Brain (VeBrain), a unified framework for perception, reasoning, and control in real world. VeBrain reformulates robotic control into common text-based MLLM tasks in the 2D visual space, thus unifying the objectives and mapping spaces of different tasks. Then, a novel robotic adapter is proposed to convert textual control signals from MLLMs to motion policies of real robots. From the data perspective, we further introduce VeBrain-600k, a high-quality instruction dataset encompassing various capabilities of VeBrain. In VeBrain-600k, we take hundreds of hours to collect, curate and annotate the data, and adopt multimodal chain-of-thought(CoT) to mix the different capabilities into a single conversation. Extensive experiments on 13 multimodal benchmarks and 5 spatial intelligence benchmarks demonstrate the superior performance of VeBrain to existing MLLMs like Qwen2.5-VL. When deployed to legged robots and robotic arms, VeBrain shows strong adaptability, flexibility, and compositional capabilities compared to existing methods. For example, compared to Qwen2.5-VL, VeBrain not only achieves substantial gains on MMVet by +5.6%, but also excels in legged robot tasks with +50% average gains.

📄 PDF Abstract BibTeX arXiv:2506.00123

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial Reasoning

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
ADOPT Please enter a description about the method here
Adapter 설명 없음

Similar Papers 제목 키워드 기반

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination

2026-07-15 · Haotian Liang, Mingkang Chen, Yufei Huang, Yuchun Guo 외 hf

Embodied cognition requires agents to connect high-level task reasoning with the physical states to be achieved. We introduce Hy-Embodied-RxBrain, an embodied cognition foundation model with joint language-visual reasoni…

Scene UnderstandingVisual ReasoningDecision Making

iFLYTEK-Embodied-Omni Technical Report

2026-06-24 · Yuan Zhang, Jingfei Ni, Guanchen Lu, Shiqi Zhang 외 arxiv

General-purpose embodied agents must understand multimodal instructions, anticipate how their environment will evolve, and produce precise control actions over extended horizons. Existing approaches typically specialize …

Video Generation

Robobench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain

2025-10-20 · Yulin Luo, Chun-Kai Fan, Menghang Dong, Jiayu Shi 외 arxiv

Building robots that can perceive, reason, and act in dynamic, unstructured environments remains a central challenge. Recent embodied systems often follow a dual-system paradigm, where System 2 performs high-level reason…

LLM as A Robotic Brain: Unifying Egocentric Memory and Control

2023-04-19 · Jinjie Mai, Jun Chen, Bing Li, Guocheng Qian 외

Embodied AI focuses on the study and development of intelligent systems that possess a physical or virtual embodiment (i.e. robots) and are able to dynamically interact with their environment. Memory and control are the …

Embodied Question AnsweringLanguage ModelingLanguage ModellingQuestion Answering+1

PaLM-E: An Embodied Multimodal Language Model

2023-03-06 · Danny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch 외

Large language models excel at a wide range of complex tasks. However, enabling general inference in the real world, e.g., for robotics problems, raises the challenge of grounding. We propose embodied language models to …

Language ModelingLanguage ModellingLarge Language Modelmodel+4