paper-with-me

Papers

Scaling may be all you need for achieving human-level object recognition capacity with human-like visual experience

2023-08-07 · A. Emin Orhan

This paper asks whether current self-supervised learning methods, if sufficiently scaled up, would be able to reach human-level visual object recognition capabilities with the same type and amount of visual experience humans learn from. Previous work on this question only considered the scaling of data size. Here, we consider the simultaneous scaling of data size, model size, and image resolution. We perform a scaling experiment with vision transformers up to 633M parameters in size (ViT-H/14) trained with up to 5K hours of human-like video data (long, continuous, mostly egocentric videos) with image resolutions of up to 476x476 pixels. The efficiency of masked autoencoders (MAEs) as a self-supervised learning algorithm makes it possible to run this scaling experiment on an unassuming academic budget. We find that it is feasible to reach human-level object recognition capacity at sub-human scales of model size, data size, and image size, if these factors are scaled up simultaneously. To give a concrete example, we estimate that a 2.5B parameter ViT model trained with 20K hours (2.3 years) of human-like video data with a spatial resolution of 952x952 pixels should be able to reach roughly human-level accuracy on ImageNet. Human-level competence is thus achievable for a fundamental perceptual capability from human-like perceptual experience (human-like in both amount and type) with extremely generic learning algorithms and architectures and without any substantive inductive biases.

📄 PDF Abstract BibTeX arXiv:2308.03712

Code (1)

eminorhan/humanlike-vits 공식 구현 pytorch

Tasks

AllObject RecognitionSelf-Supervised Learning

Similar Papers 제목 키워드 기반

How much human-like visual experience do current self-supervised learning algorithms need in order to achieve human-level object recognition?

2021-09-23 · A. Emin Orhan

This paper addresses a fundamental question: how good are our current self-supervised visual representation learning algorithms relative to humans? More concretely, how much "human-like" natural visual experience would t…

Object RecognitionRepresentation LearningSelf-Supervised Learning

Can Language Models Understand Physical Concepts?

2023-05-23 · Lei LI, Jingjing Xu, Qingxiu Dong, Ce Zheng 외

Language models~(LMs) gradually become general-purpose interfaces in the interactive and embodied world, where the understanding of physical concepts is an essential prerequisite. However, it is not yet clear whether LMs…

Scaling Laws for Task-Optimized Models of the Primate Visual Ventral Stream

2024-11-08 · Abdulkadir Gokce, Martin Schrimpf

When trained on large-scale object classification datasets, certain artificial neural network models begin to approximate core object recognition (COR) behaviors and neural response patterns in the primate visual ventral…

Brain DecodingInductive BiasObject Recognition

Taccel: Scaling Up Vision-based Tactile Robotics via High-performance GPU Simulation

2025-04-17 · Yuyang Li, Wenxin Du, Chang Yu, Puhao Li 외

Tactile sensing is crucial for achieving human-level robotic capabilities in manipulation tasks. VBTSs have emerged as a promising solution, offering high spatial resolution and cost-effectiveness by sensing contact thro…

GPUObject RecognitionRobotic Grasping

Joint Learning of Pose Regression and Denoising Diffusion with Score Scaling Sampling for Category-level 6D Pose Estimation

2025-10-05 · Seunghyun Lee, Tae-Kyun Kim arxiv

Latest diffusion models have shown promising results in category-level 6D object pose estimation by modeling the conditional pose distribution with depth image input. The existing methods, however, suffer from slow conve…

6D Pose Estimation