paper-with-me

홈 › Papers

What Makes Pre-Trained Visual Representations Successful for Robust Manipulation?

2023-11-03 · Kaylee Burns, Zach Witzel, Jubayer Ibn Hamid, Tianhe Yu, Chelsea Finn, Karol Hausman

Inspired by the success of transfer learning in computer vision, roboticists have investigated visual pre-training as a means to improve the learning efficiency and generalization ability of policies learned from pixels. To that end, past work has favored large object interaction datasets, such as first-person videos of humans completing diverse tasks, in pursuit of manipulation-relevant features. Although this approach improves the efficiency of policy learning, it remains unclear how reliable these representations are in the presence of distribution shifts that arise commonly in robotic applications. Surprisingly, we find that visual representations designed for manipulation and control tasks do not necessarily generalize under subtle changes in lighting and scene texture or the introduction of distractor objects. To understand what properties do lead to robust representations, we compare the performance of 15 pre-trained vision models under different visual appearances. We find that emergent segmentation ability is a strong predictor of out-of-distribution generalization among ViT models. The rank order induced by this metric is more predictive than metrics that have previously guided generalization research within computer vision and machine learning, such as downstream ImageNet accuracy, in-domain accuracy, or shape-bias as evaluated by cue-conflict performance. We test this finding extensively on a suite of distribution shifts in ten tasks across two simulated manipulation environments. On the ALOHA setup, segmentation score predicts real-world performance after offline training with 50 demonstrations.

📄 PDF Abstract BibTeX arXiv:2312.12444

Code (0)

등록된 구현이 없습니다.

Tasks

Out-of-Distribution GeneralizationTransfer Learning

Similar Papers 제목 키워드 기반

Look, Listen and Learn

2017-05-23 · ICCV 2017 10 · Relja Arandjelović, Andrew Zisserman

We consider the question: what can be learnt by looking at and listening to a large number of unlabelled videos? There is a valuable, but so far untapped, source of information contained in the video itself -- the corres…

Audio ClassificationGeneral ClassificationSound Classification

Visualising Deep Network's Time-Series Representations

2021-03-12 · Błażej Leporowski, Alexandros Iosifidis

Despite the popularisation of machine learning models, more often than not, they still operate as black boxes with no insight into what is happening inside the model. There exist a few methods that allow to visualise and…

Time SeriesTime Series AnalysisTime Series Classification

Towards robust vision by multi-task learning on monkey visual cortex

2021-07-29 · NeurIPS 2021 12 · Shahd Safarani, Arne Nix, Konstantin Willeke, Santiago A. Cadena 외

Deep neural networks set the state-of-the-art across many tasks in computer vision, but their generalization ability to image distortions is surprisingly fragile. In contrast, the mammalian visual system is robust to a w…

image-classificationImage ClassificationMulti-Task LearningOut-of-Distribution Generalization

What You Say Is What You Show: Visual Narration Detection in Instructional Videos

2023-01-05 · Kumar Ashutosh, Rohit Girdhar, Lorenzo Torresani, Kristen Grauman

Narrated ''how-to'' videos have emerged as a promising data source for a wide range of learning problems, from learning visual representations to training robot policies. However, this data is extremely noisy, as the nar…

Subject-Aware Multi-Granularity Alignment for Zero-Shot EEG-to-Image Retrieval

2026-04-20 · Lin Jiang, Qingshan She, Jiale Xu, Haiqi Xu 외 arxiv

Decoding visual content from electroencephalography (EEG) is important for understanding neural visual representations and developing non-invasive brain-computer interfaces. Existing approaches mainly improve EEG represe…

Image Retrieval