paper-with-me

Papers

Beyond Viewpoint Generalization: What Multi-View Demonstrations Offer and How to Synthesize Them for Robot Manipulation?

2026-03-23 · Boyang Cai, Qiwei Liang, Jiawei Li, Shihang Weng, Zhaoxin Zhang, Tao Lin, Xiangyu Chen, Wenjie Zhang, Jiaqi Mao, Weisheng Xu, Bin Yang, Jiaming Liang, Junhao Cai, Renjing Xu arxiv

Does multi-view demonstration truly improve robot manipulation, or merely enhance cross-view robustness? We present a systematic study quantifying the performance gains, scaling behavior, and underlying mechanisms of multi-view data for robot manipulation. Controlled experiments show that, under both fixed and randomized backgrounds, multi-view demonstrations consistently improve single-view policy success and generalization. Performance varies non-monotonically with view coverage, revealing effective regimes rather than a simple "more is better" trend. Notably, multi-view data breaks the scaling limitation of single-view datasets and continues to raise performance ceilings after saturation. Mechanistic analysis shows that multi-view learning promotes manipulation-relevant visual representations, better aligns the action head with the learned feature distribution, and reduces overfitting. Motivated by the importance of multi-view data and its scarcity in large-scale robotic datasets, as well as the difficulty of collecting additional viewpoints in real world settings, we propose RoboNVS, a geometry-aware self-supervised framework that synthesizes novel-view videos from monocular inputs. The generated data consistently improves downstream policies in both simulation and real-world environments.

📄 PDF Abstract BibTeX arXiv:2603.26757

Code (0)

등록된 구현이 없습니다.

Tasks

Robot Manipulation

Similar Papers 제목 키워드 기반

On the Capability of CNNs to Generalize to Unseen Category-Viewpoint Combinations

2021-01-01 · Spandan Madan, Timothy Henry, Jamell Arthur Dozier, Helen Ho 외

Object recognition and viewpoint estimation lie at the heart of visual understanding. Recent works suggest that convolutional neural networks (CNNs) fail to generalize to category-viewpoint combinations not seen during t…

Object RecognitionViewpoint Estimation

Predicting Camera Viewpoint Improves Cross-dataset Generalization for 3D Human Pose Estimation

2020-04-07 · Zhe Wang, Daeyun Shin, Charless C. Fowlkes

Monocular estimation of 3d human pose has attracted increased attention with the availability of large ground-truth motion capture datasets. However, the diversity of training data available is limited and it is not clea…

3D Human Pose EstimationDiversityMonocular 3D Human Pose EstimationPose Estimation

Look-Before-Move: Narrative-Grounded World Visual Attention in Dynamic 3D Story Worlds

2026-06-25 · Jiaming Bian, Bingliang Li, Yuehao Wu, Pichao Wang 외 arxiv

As embodied AI and world models increasingly operate in dynamic 3D environments, visual perception must move beyond passively interpreting given observations toward actively deciding what to observe. We study this proble…

3D Neural Scene Representations for Visuomotor Control

2021-07-08 · Yunzhu Li, Shuang Li, Vincent Sitzmann, Pulkit Agrawal 외

Humans have a strong intuitive understanding of the 3D environment around us. The mental model of the physics in our brain applies to objects of different materials and enables us to perform a wide range of manipulation …

Contrastive LearningFuture predictionNeRFNovel View Synthesis

Pose-Aware Self-Supervised Learning with Viewpoint Trajectory Regularization

2024-03-22 · Jiayun Wang, Yubei Chen, Stella X. Yu

Learning visual features from unlabeled images has proven successful for semantic categorization, often by mapping different $views$ of the same object to the same feature to achieve recognition invariance. However, visu…

ObjectPose EstimationRepresentation LearningSelf-Supervised Learning