paper-with-me

홈 › Papers

Training Vision Transformers with Only 2040 Images

2022-01-26 · Yun-Hao Cao, Hao Yu, Jianxin Wu

Vision Transformers (ViTs) is emerging as an alternative to convolutional neural networks (CNNs) for visual recognition. They achieve competitive results with CNNs but the lack of the typical convolutional inductive bias makes them more data-hungry than common CNNs. They are often pretrained on JFT-300M or at least ImageNet and few works study training ViTs with limited data. In this paper, we investigate how to train ViTs with limited data (e.g., 2040 images). We give theoretical analyses that our method (based on parametric instance discrimination) is superior to other methods in that it can capture both feature alignment and instance similarities. We achieve state-of-the-art results when training from scratch on 7 small datasets under various ViT backbones. We also investigate the transferring ability of small datasets and find that representations learned from small datasets can even improve large-scale ImageNet training.

📄 PDF Abstract BibTeX arXiv:2201.10728

Code (2)

CupidJay/Training-Vision-Transformers-with-only-2040-images 공식 구현 pytorch
niranjankrishna-acad/Training-Vision-Transformers-with-Only-2040-Images pytorch

Tasks

Inductive Bias

Similar Papers 제목 키워드 기반

Pre-training Vision Transformers with Very Limited Synthesized Images

2023-07-27 · ICCV 2023 1 · Ryo Nakamura, Hirokatsu Kataoka, Sora Takashima, Edgar Josafat Martinez Noriega 외

Formula-driven supervised learning (FDSL) is a pre-training method that relies on synthetic images generated from mathematical formulae such as fractals. Prior work on FDSL has shown that pre-training vision transformers…

Data Augmentation

Self-supervised Vision Transformers for 3D Pose Estimation of Novel Objects

2023-05-31 · Stefan Thalhammer, Jean-Baptiste Weibel, Markus Vincze, Jose Garcia-Rodriguez

Object pose estimation is important for object manipulation and scene understanding. In order to improve the general applicability of pose estimators, recent research focuses on providing estimates for novel objects, tha…

3D Pose EstimationContrastive LearningObjectPose Estimation+3

Learning Explicit Object-Centric Representations with Vision Transformers

2022-10-25 · Oscar Vikström, Alexander Ilin

With the recent successful adaptation of transformers to the vision domain, particularly when trained in a self-supervised fashion, it has been shown that vision transformers can learn impressive object-reasoning-like be…

ObjectSegmentationSemantic Segmentation

RGB no more: Minimally-decoded JPEG Vision Transformers

2022-11-29 · CVPR 2023 1 · Jeongsoo Park, Justin Johnson

Most neural networks for computer vision are designed to infer using RGB images. However, these RGB images are commonly encoded in JPEG before saving to disk; decoding them imposes an unavoidable overhead for RGB network…

Data Augmentation

Machine Learning for Brain Disorders: Transformers and Visual Transformers

2023-03-21 · Robin Courant, Maika Edberg, Nicolas Dufour, Vicky Kalogeiton

Transformers were initially introduced for natural language processing (NLP) tasks, but fast they were adopted by most deep learning fields, including computer vision. They measure the relationships between pairs of inpu…

Decoderimage-classificationImage Classification