paper-with-me

Papers

How much human-like visual experience do current self-supervised learning algorithms need in order to achieve human-level object recognition?

2021-09-23 · A. Emin Orhan

This paper addresses a fundamental question: how good are our current self-supervised visual representation learning algorithms relative to humans? More concretely, how much "human-like" natural visual experience would these algorithms need in order to reach human-level performance in a complex, realistic visual object recognition task such as ImageNet? Using a scaling experiment, here we estimate that the answer is several orders of magnitude longer than a human lifetime: typically on the order of a million to a billion years of natural visual experience (depending on the algorithm used). We obtain even larger estimates for achieving human-level performance in ImageNet-derived robustness benchmarks. The exact values of these estimates are sensitive to some underlying assumptions, however even in the most optimistic scenarios they remain orders of magnitude larger than a human lifetime. We discuss the main caveats surrounding our estimates and the implications of these surprising results.

📄 PDF Abstract BibTeX arXiv:2109.11523

Code (1)

eminorhan/human-ssl 공식 구현 pytorch

Tasks

Object RecognitionRepresentation LearningSelf-Supervised Learning

Similar Papers 제목 키워드 기반

Visual Memorability for Robotic Interestingness via Unsupervised Online Learning

2020-05-18 · ECCV 2020 8 · Chen Wang, Wenshan Wang, Yuheng Qiu, Yafei Hu 외

In this paper, we explore the problem of interesting scene prediction for mobile robots. This area is currently underexplored but is crucial for many practical applications such as autonomous exploration and decision mak…

Decision MakingIncremental LearningScene RecognitionTranslation

Scaling may be all you need for achieving human-level object recognition capacity with human-like visual experience

2023-08-07 · A. Emin Orhan

This paper asks whether current self-supervised learning methods, if sufficiently scaled up, would be able to reach human-level visual object recognition capabilities with the same type and amount of visual experience hu…

AllObject RecognitionSelf-Supervised Learning

The BabyView dataset: High-resolution egocentric videos of infants' and young children's everyday experiences

2024-06-14 · Bria Long, Violet Xiang, Stefan Stojanov, Robert Z. Sparks 외

Human children far exceed modern machine learning algorithms in their sample efficiency, achieving high performance in key domains with much less data than current models. This ''data gap'' is a key challenge both for bu…

Depth EstimationImage SegmentationObject RecognitionPose Estimation+3

Temporal Slowness in Central Vision Drives Semantic Object Learning

2026-02-04 · Timothy Schaumlöffel, Arthur Aubret, Gemma Roig, Jochen Triesch arxiv

Humans acquire semantic object representations from egocentric visual streams with minimal supervision, but the underlying mechanisms remain unclear. Importantly, the visual system only processes the center of its field …

Self-Supervised Learning

Visual Prediction of Priors for Articulated Object Interaction

2020-06-06 · Caris Moses, Michael Noseworthy, Leslie Pack Kaelbling, Tomás Lozano-Pérez 외

Exploration in novel settings can be challenging without prior experience in similar domains. However, humans are able to build on prior experience quickly and efficiently. Children exhibit this behavior when playing wit…

ObjectPrediction