paper-with-me

홈 › Papers

Act, Sense, Act: Learning Active Perception from Large-Scale Egocentric Human Data

2026-02-04 · Jialiang Li, Yi Qiao, Yunhan Guo, Changwen Chen, Wenzhao Lian arxiv

Achieving generalizable manipulation in unconstrained environments requires the robot to proactively resolve information uncertainty, i.e., the capability of active perception. However, existing methods are often confined in limited types of sensing behaviors, restricting their applicability to complex environments. In this work, we formalize active perception as a history-dependent perception-action loop driven by information-seeking action and decision branching, providing a structured categorization of visual active perception paradigms. Building on this perspective, we introduce CoMe-VLA, a cognitive and memory-aware vision-language-action (VLA) framework that leverages large-scale human egocentric data to learn versatile exploration and manipulation priors. Our framework integrates a cognitive auxiliary head for autonomous sub-task transitions and a dual-track memory system to maintain consistent self and environmental awareness by fusing proprioceptive and visual temporal contexts. By aligning human and robot hand-eye coordination behaviors in a unified egocentric action space, we train the model progressively in three stages. Extensive experiments on a wheel-based humanoid have demonstrated strong robustness and adaptability of our proposed method across diverse long-horizon tasks spanning multiple active perception scenarios.

📄 PDF Abstract BibTeX arXiv:2602.04600

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RoboSense: Large-scale Dataset and Benchmark for Egocentric Robot Perception and Navigation in Crowded and Unstructured Environments

2024-08-28 · CVPR 2025 1 · Haisheng Su, Feixiang Song, Cong Ma, Wei Wu 외

Reliable embodied perception from an egocentric perspective is challenging yet essential for autonomous navigation technology of intelligent mobile agents. With the growing demand of social robotics, near-field scene und…

Autonomous DrivingAutonomous Navigationobject-detectionObject Detection+3

ActiveMimic: Egocentric Video Pretraining with Active Perception

2026-06-04 · Xingyao Lin, Guojin Zhong, Tianyi Lu, Ziyi Ye 외 arxiv

Egocentric human video offers a scalable alternative to robot data for pretraining, yet models pretrained on such video consistently underperform those pretrained on robot data. We attribute this gap to a missing signal,…

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation

2026-07-16 · Chi Kit Wong, Ye Pan, Yuanhuiyi Lyu, Xu Zheng 외 arxiv

Egocentric Visual Question Answering (VQA) has attracted widespread attention as an important task for enabling Multimodal Large Language Models (MLLMs) to interact with the real world. However, existing MLLMs struggle t…

Visual Question AnsweringSpatial Reasoning

Behavior Cloning for Active Perception with Low-Resolution Egocentric Vision

2026-05-13 · Anthony Bilic, Chen Chen, Ladislau Bölöni arxiv

We investigate whether behavior cloning is sufficient to produce active perception in a structured object-finding task. A low-cost robot arm equipped with a wrist-mounted egocentric RGB camera must reposition to center a…

DreamNav: A Trajectory-Based Imaginative Framework for Zero-Shot Vision-and-Language Navigation

2025-09-14 · Yunheng Wang, Yuetong Fang, Taowen Wang, Yixiao Feng 외 arxiv

Vision-and-Language Navigation in Continuous Environments (VLN-CE), which links language instructions to perception and control in the real world, is a core capability of embodied robots. Recently, large-scale pretrained…

Scene Understanding