paper-with-me

홈 › Papers

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination

2026-07-15 · Haotian Liang, Mingkang Chen, Yufei Huang, Yuchun Guo, Xiaomeng Zhu, Xiangli Shi, Kaixuan Wang, Yunxuan Mao, Weijie Zhou, Ling Chen, Shirong Zeng, Yueyu Long, Yuchen Si, Yajuan Zhu, Xingyu Zhou, Minghui Wang, Wanjia He, Xin Yang, Lingzhu Xiang, Zhiqing Liu, Bohan Ma, Xiran Huang, Tianshuo Yang, Zhiheng Liu, Xuantang Xiong, Zisheng Lu, Ping Luo, Yao Mu, Han Hu, Zhengyou Zhang hf

Embodied cognition requires agents to connect high-level task reasoning with the physical states to be achieved. We introduce Hy-Embodied-RxBrain, an embodied cognition foundation model with joint language-visual reasoning and imagination. Unlike vision-language models that emphasize scene understanding and textual decision making, or generative world models that mainly predict future visual states, RxBrain represents embodied plans in a single planning sequence where language and visual imagination play complementary roles. Language provides the abstract structure of a plan, including task decomposition, planning primitives, constraints, temporal order, and decision logic, while visual imagination grounds this structure through world state prediction and joint subgoal planning, associating each planning step with intermediate and final physical states. RxBrain adopts a unified multimodal Mixture-of-Transformers architecture that supports language, image, and video understanding and generation within one model. To train this capability, we build an automatic pipeline that converts embodied videos into joint text-visual planning supervision by decomposing videos into planning steps and aligning them with visual state transitions. We further introduce RxBrain-Bench to evaluate whether models can represent embodied plans through joint textual and visual components rather than separate understanding or generation. Experiments show that RxBrain maintains embodied understanding and generation abilities, and produces plans with coupled textual reasoning, world state prediction, and joint subgoal planning. We also extend RxBrain to continuous robot action generation, where it shows promising real-robot performance without large-scale action-data pretraining. These results provide an initial step toward foundation models for embodied cognition.

📄 PDF Abstract BibTeX arXiv:2607.14187

Code (0)

등록된 구현이 없습니다.

Tasks

Scene UnderstandingVisual ReasoningDecision Making

Similar Papers 제목 키워드 기반

Towards Embodied Cognition in Robots via Spatially Grounded Synthetic Worlds

2025-05-20 · Joel Currie, Gioele Migno, Enrico Piacenti, Maria Elena Giannaccini 외

We present a conceptual framework for training Vision-Language Models (VLMs) to perform Visual Perspective Taking (VPT), a core capability for embodied cognition essential for Human-Robot Interaction (HRI). As a first st…

Spatial Reasoning

ECBench: Can Multi-modal Foundation Models Understand the Egocentric World? A Holistic Embodied Cognition Benchmark

2025-01-09 · CVPR 2025 1 · Ronghao Dang, Yuqian Yuan, Wenqi Zhang, Yifei Xin 외

The enhancement of generalization in robots by large vision-language models (LVLMs) is increasingly evident. Therefore, the embodied cognitive abilities of LVLMs based on egocentric videos are of great interest. However,…

FairnessHallucinationQuestion AnsweringVideo Question Answering

RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model

2026-07-20 · Kehan Li, Bohan Hou, Minghao Zhu, Tianyi Zhang 외 hf

We present RynnBrain 1.1, a family of embodied foundation models spanning 2B, 9B, and 122B-A10B scales. Trained with a unified spatio-temporal and physically grounded framework, RynnBrain 1.1 supports embodied perception…

Robot ManipulationSpatial Reasoning

RynnEC: Bringing MLLMs into Embodied World

2025-08-19 · Ronghao Dang, Yuqian Yuan, Yunxuan Mao, Kehan Li 외 arxiv

We introduce RynnEC, a video multimodal large language model designed for embodied cognition. Built upon a general-purpose vision-language foundation model, RynnEC incorporates a region encoder and a mask decoder, enabli…

Object SegmentationSpatial Reasoning

ET-Plan-Bench: Embodied Task-level Planning Benchmark Towards Spatial-Temporal Cognition with Foundation Models

2024-10-02 · Lingfeng Zhang, Yuening Wang, Hongjian Gu, Atia Hamidizadeh 외

Recent advancements in Large Language Models (LLMs) have spurred numerous attempts to apply these technologies to embodied tasks, particularly focusing on high-level task planning and task decomposition. To further explo…

DiagnosticTask Planning