paper-with-me

Papers

Thinker: A vision-language foundation model for embodied intelligence

2026-01-29 · Baiyu Pan, Daqin Luo, Junpeng Yang, Jiyuan Wang, Yixuan Zhang, Hailin Shi, Jichao Jiao arxiv

When large vision-language models are applied to the field of robotics, they encounter problems that are simple for humans yet error-prone for models. Such issues include confusion between third-person and first-person perspectives and a tendency to overlook information in video endings during temporal reasoning. To address these challenges, we propose Thinker, a large vision-language foundation model designed for embodied intelligence. We tackle the aforementioned issues from two perspectives. Firstly, we construct a large-scale dataset tailored for robotic perception and reasoning, encompassing ego-view videos, visual grounding, spatial understanding, and chain-of-thought data. Secondly, we introduce a simple yet effective approach that substantially enhances the model's capacity for video comprehension by jointly incorporating key frames and full video sequences as inputs. Our model achieves state-of-the-art results on two of the most commonly used benchmark datasets in the field of task planning.

📄 PDF Abstract BibTeX arXiv:2601.21199

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Grounding

Similar Papers 제목 키워드 기반

VLA-Thinker: Boosting Vision-Language-Action Models through Thinking-with-Image Reasoning

2026-03-15 · Chaoyang Wang, Wenrui Bao, Sicheng Gao, Bingxin Xu 외 arxiv

Vision-Language-Action (VLA) models have shown promising capabilities for embodied intelligence, but most existing approaches rely on text-based chain-of-thought reasoning where visual inputs are treated as static contex…

Reinforcement Learning

TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation

2026-06-10 · Kailin Lyu, Di Wu, Pengwei Zhang, Yuhang Zheng 외 arxiv

Touch is a key modality for embodied agents to understand the physical world. Although recent work has incorporated tactile signals into language systems for tactile commonsense reasoning, scaling such systems to realist…

Thinking with Geometry: Active Geometry Integration for Spatial Reasoning

2026-02-05 · Haoyuan Li, Qihang Cao, Tao Tang, Kun Xiang 외 arxiv

Recent progress in spatial reasoning with Multimodal Large Language Models (MLLMs) increasingly leverages geometric priors from 3D encoders. However, most existing integration strategies remain passive: geometry is expos…

Autonomous DrivingSpatial Reasoning

Survey of Vision-Language-Action Models for Embodied Manipulation

2025-08-21 · Haoran Li, Yuhui Chen, Wenbo Cui, Weiheng Liu 외 arxiv

Embodied intelligence systems, which enhance agent capabilities through continuous environment interactions, have garnered significant attention from both academia and industry. Vision-Language-Action models, inspired by…

RynnBrain: Open Embodied Foundation Models

2026-02-13 · Ronghao Dang, Jiayan Guo, Bohan Hou, Sicong Leng 외 arxiv

Despite rapid progress in multimodal foundation models, embodied intelligence community still lacks a unified, physically grounded foundation model that integrates perception, reasoning, and planning within real-world sp…

Spatial Reasoning