paper-with-me

홈 › Papers

Vision Transformers in 2022: An Update on Tiny ImageNet

2022-05-21 · Ethan Huynh

The recent advances in image transformers have shown impressive results and have largely closed the gap between traditional CNN architectures. The standard procedure is to train on large datasets like ImageNet-21k and then finetune on ImageNet-1k. After finetuning, researches will often consider the transfer learning performance on smaller datasets such as CIFAR-10/100 but have left out Tiny ImageNet. This paper offers an update on vision transformers' performance on Tiny ImageNet. I include Vision Transformer (ViT) , Data Efficient Image Transformer (DeiT), Class Attention in Image Transformer (CaiT), and Swin Transformers. In addition, Swin Transformers beats the current state-of-the-art result with a validation accuracy of 91.35%. Code is available here: https://github.com/ehuynh1106/TinyImageNet-Transformers

📄 PDF Abstract BibTeX arXiv:2205.10660

Code (1)

ehuynh1106/TinyImageNet-Transformers 공식 구현 pytorch

Tasks

Image ClassificationTransfer Learning

Methods 이 논문이 사용한 방법론

Vision Transformer The Vision Transformer, or ViT, is a model for image classification that employs a Transformer-like architecture over…

Similar Papers 제목 키워드 기반

Parameter Efficient Continual Learning for Sparse Event-Based Transformers

2026-08-27 · Vaishnavi Nagabhushana, Kartikay Agrawal, Ayon Borthakur arxiv

Robotic and edge intelligence systems operate in dynamic environments where data arrives continuously, requiring models to adapt while preserving previously learned knowledge under strict memory and energy constraints. W…

parameter-efficient fine-tuningclass-incremental learningContinual LearningEvent-based vision

Revisiting Residual Connections: Orthogonal Updates for Stable and Efficient Deep Networks

2025-05-17 · Giyeong Oh, Woohyun Cho, Siyeol Kim, Suhwan Choi 외

Residual connections are pivotal for deep neural networks, enabling greater depth by mitigating vanishing gradients. However, in standard residual updates, the module's output is directly added to the input stream. This …

TinyViT: Fast Pretraining Distillation for Small Vision Transformers

2022-07-21 · Kan Wu, Jinnian Zhang, Houwen Peng, Mengchen Liu 외

Vision transformer (ViT) recently has drawn great attention in computer vision due to its remarkable model capability. However, most prevailing ViT models suffer from huge number of parameters, restricting their applicab…

Image ClassificationKnowledge Distillation

CNN and ViT Efficiency Study on Tiny ImageNet and DermaMNIST Datasets

2025-05-13 · Aidar Amangeldi, Angsar Taigonyrov, Muhammad Huzaid Jawad, Chinedu Emmanuel Mbonu

This study evaluates the trade-offs between convolutional and transformer-based architectures on both medical and general-purpose image classification benchmarks. We use ResNet-18 as our baseline and introduce a fine-tun…

image-classificationImage Classification

GenFormer -- Generated Images are All You Need to Improve Robustness of Transformers on Small Datasets

2024-08-26 · Sven Oehri, Nikolas Ebert, Ahmed Abdullah, Didier Stricker 외

Recent studies showcase the competitive accuracy of Vision Transformers (ViTs) in relation to Convolutional Neural Networks (CNNs), along with their remarkable robustness. However, ViTs demand a large amount of data to a…

AllData Augmentationimage-classificationImage Classification+1