paper-with-me

Papers

SaPaVe: Towards Active Perception and Manipulation in Vision-Language-Action Models for Robotics

2026-03-12 · Mengzhen Liu, Enshen Zhou, Cheng Chi, Yi Han, Shanyu Rong, Liming Chen, Pengwei Wang, Zhongyuan Wang, Shanghang Zhang arxiv

Active perception and manipulation are crucial for robots to interact with complex scenes. Existing methods struggle to unify semantic-driven active perception with robust, viewpoint-invariant execution. We propose SaPaVe, an end-to-end framework that jointly learns these capabilities in a data-efficient manner. Our approach decouples camera and manipulation actions rather than placing them in a shared action space, and follows a bottom-up training strategy: we first train semantic camera control on a large-scale dataset, then jointly optimize both action types using hybrid data. To support this framework, we introduce ActiveViewPose-200K, a dataset of 200k image-language-camera movement pairs for semantic camera movement learning, and a 3D geometry-aware module that improves execution robustness under dynamic viewpoints. We also present ActiveManip-Bench, the first benchmark for evaluating active manipulation beyond fixed-view settings. Extensive experiments in both simulation and real-world environments show that SaPaVe outperforms recent vision-language-action models such as GR00T N1 and \(π_0\), achieving up to 31.25\% higher success rates in real-world tasks. These results show that tightly coupled perception and execution, when trained with decoupled yet coordinated strategies, enable efficient and generalizable active manipulation. Project page: https://lmzpai.github.io/SaPaVe

📄 PDF Abstract BibTeX arXiv:2603.12193

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic Manipulation

2026-01-13 · Zhenyang Liu, Yongchong Gu, Yikai Wang, Xiangyang Xue 외 arxiv

Recent advances in robot manipulation have leveraged pre-trained vision-language models (VLMs) and explored integrating 3D spatial signals into these models for effective action prediction, giving rise to the promising v…

Robot Manipulation

ActiveGlasses: Learning Manipulation with Active Vision from Ego-centric Human Demonstration

2026-04-09 · Yanwen Zou, Chenyang Shi, Wenye Yu, Han Xue 외 arxiv

Large-scale real-world robot data collection is a prerequisite for bringing robots into everyday deployment. However, existing pipelines often rely on specialized handheld devices to bridge the embodiment gap, which not …

Robot Manipulation

Act, Sense, Act: Learning Active Perception from Large-Scale Egocentric Human Data

2026-02-04 · Jialiang Li, Yi Qiao, Yunhan Guo, Changwen Chen 외 arxiv

Achieving generalizable manipulation in unconstrained environments requires the robot to proactively resolve information uncertainty, i.e., the capability of active perception. However, existing methods are often confine…

Towards Exploratory and Focused Manipulation with Bimanual Active Perception: A New Problem, Benchmark and Strategy

2026-02-02 · Yuxin He, Ruihao Zhang, Tianao Shen, Cheng Liu 외 arxiv

Recently, active vision has reemerged as an important concept for manipulation, since visual occlusion occurs more frequently when main cameras are mounted on the robot heads. We reflect on the visual occlusion issue and…

I-Perceive: A Foundation Model for Active Perception with Language Instructions

2026-02-28 · Yongxi Huang, Zhuohang Wang, Wenjing Tang, Cewu Lu 외 arxiv

Active perception, the ability of a robot to proactively adjust its viewpoint to acquire task-relevant information, is essential for robust operation in unstructured real-world environments. While critical for downstream…

Zero-shot GeneralizationInstruction Following