paper-with-me

Papers

MV-Actor: Aligning Multi-View Semantics and Spatial Awareness for Bimanual Manipulation

2026-06-09 · Yinchen Tian, Huan Li, Muyao Peng, Xi Wang, Yan Wang, You Yang arxiv

Robotic manipulation has been widely applied in industrial scenarios. Compared with single-arm manipulation, bimanual manipulation is equipped with multiple cameras to capture information from different viewpoints. However, existing multi-view policies encode each view independently or fuse view features shallowly, resulting in limited sharing semantic perception and unreliable spatial awareness. In this paper, we propose \textbf{MV-Actor}, a multi-view perception framework that builds a unified semantic-spatial representation for bimanual manipulation. First, MV-Actor performs Multi-view Semantic Interaction to share semantic perception across views. Then it uses Semantic-Spatial Token Interaction to ground visual semantics with feed-forward reconstruction model features and acquire reliable spatial awareness. Finally, a Guided Metric Depth Repair module refines degraded sensor depth to provide more reliable metric anchors under consumer-grade depth noise. In simulation experiments conducted on the PerAct2 bimanual benchmark, MV-Actor achieves a state-of-the-art average success rate of 87.8\%. In real-world evaluations with more frequent viewpoint changes and unstable consumer-grade depth, MV-Actor outperforms both RGB and RGB-D baselines, further demonstrating the benefit of sharing semantic perception and reliable spatial awareness for bimanual manipulation.

📄 PDF Abstract BibTeX arXiv:2606.10899

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Learning 3D Representations for Spatial Intelligence from Unposed Multi-View Images

2026-04-12 · Bo Zhou, Qiuxia Lai, Zeren Sun, Xiangbo Shu 외 arxiv

Robust 3D representation learning forms the perceptual foundation of spatial intelligence, enabling downstream tasks in scene understanding and embodied AI. However, learning such representations directly from unposed mu…

Representation LearningScene Understanding

SpatialActor: Exploring Disentangled Spatial Representations for Robust Robotic Manipulation

2025-11-12 · Hao Shi, Bin Xie, Yingfei Liu, Yang Yue 외 arxiv

Robotic manipulation requires precise spatial understanding to interact with objects in the real world. Point-based methods suffer from sparse sampling, leading to the loss of fine-grained semantics. Image-based methods …

What Does the Brain See? Multiview Neural Representations to Demystify the Brain-Visual Alignment

2026-06-24 · Salini Yadav, Taveena Lotey, Pravendra Singh, Partha Pratim Roy arxiv

Zero-shot visual decoding from electroencephalography (EEG) aims to infer visual semantics from non-invasive neural recordings, but remains challenging due to the low signal-to-noise ratio, non-stationarity, and limited …

Representation LearningContrastive LearningGraph Learning

m2sv: A Scalable Benchmark for Map-to-Street-View Spatial Reasoning

2026-01-27 · Yosub Shin, Michael Buriek, Igor Molybog arxiv

Vision--language models (VLMs) achieve strong performance on many multimodal benchmarks but remain brittle on spatial reasoning tasks that require aligning abstract overhead representations with egocentric views. We intr…

Reinforcement LearningSpatial Reasoning

Imputation-free and Alignment-free: Incomplete Multi-view Clustering Driven by Consensus Semantic Learning

2025-05-16 · CVPR 2025 1 · Yuzhuo Dai, Jiaqi Jin, Zhibin Dong, Siwei Wang 외

In incomplete multi-view clustering (IMVC), missing data induce prototype shifts within views and semantic inconsistencies across views. A feasible solution is to explore cross-view consistency in paired complete observa…

ClusteringGraph ClusteringImputationIncomplete multi-view clustering