paper-with-me

Papers

I-Perceive: A Foundation Model for Active Perception with Language Instructions

2026-02-28 · Yongxi Huang, Zhuohang Wang, Wenjing Tang, Cewu Lu, Panpan Cai arxiv

Active perception, the ability of a robot to proactively adjust its viewpoint to acquire task-relevant information, is essential for robust operation in unstructured real-world environments. While critical for downstream tasks such as manipulation, existing approaches have largely been confined to local settings (e.g., table-top scenes) with fixed perception objectives (e.g., occlusion reduction). Addressing active perception with open-ended intents in large-scale environments remains an open challenge. To bridge this gap, we propose I-Perceive, a foundation model for active perception conditioned on natural language instructions, designed for mobile manipulators and indoor environments. I-Perceive predicts camera views that follows open-ended language instructions, based on image-based scene contexts. By fusing a Vision-Language Model (VLM) backbone with a geometric foundation model, I-Perceive bridges semantic and geometric understanding, thus enabling effective reasoning for active perception. We train I-Perceive on a diverse dataset comprising real-world scene-scanning data and simulation data, both processed via an automated and scalable data generation pipeline. Experiments demonstrate that I-Perceive significantly outperforms state-of-the-art VLMs in both prediction accuracy and instruction following of generated camera views, and exhibits strong zero-shot generalization to novel scenes and tasks.

📄 PDF Abstract BibTeX arXiv:2603.00600

Code (0)

등록된 구현이 없습니다.

Tasks

Zero-shot GeneralizationInstruction Following

Similar Papers 제목 키워드 기반

Embodied Spatial Intelligence: from Implicit Scene Modeling to Spatial Reasoning

2025-08-30 · Jiading Fang arxiv

This thesis introduces "Embodied Spatial Intelligence" to address the challenge of creating robots that can perceive and act in the real world based on natural language instructions. To bridge the gap between Large Langu…

Spatial Reasoning

ProFocus: Proactive Perception and Focused Reasoning in Vision-and-Language Navigation

2026-03-01 · Wei Xue, Mingcheng Li, Xuecheng Wu, Jingqun Tang 외 arxiv

Vision-and-Language Navigation (VLN) requires agents to accurately perceive complex visual environments and reason over navigation instructions and histories. However, existing methods passively process redundant visual …

ExploreVLM: Closed-Loop Robot Exploration Task Planning with Vision-Language Models

2025-08-16 · Zhichen Lou, Kechun Xu, Zhongxiang Zhou, Rong Xiong arxiv

The advancement of embodied intelligence is accelerating the integration of robots into daily life as human assistants. This evolution requires robots to not only interpret high-level instructions and plan tasks but also…

From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models

2025-09-29 · Chenyue Zhou, Mingxuan Wang, Yanbiao Ma, Chenxu Wu 외 arxiv

Multimodal Large Language Models (MLLMs) strive to achieve a profound, human-like understanding of and interaction with the physical world, but often exhibit a shallow and incoherent integration when acquiring informatio…

Teaching Perception

2019-11-21 · Jonathan Connell

The visual world is very rich and generally too complex to perceive in its entirety. Yet only certain features are typically required to adequately perform some task in a given situation. Rather than hardwire-in decision…