paper-with-me

홈 › Papers

CNN and ViT Efficiency Study on Tiny ImageNet and DermaMNIST Datasets

2025-05-13 · Aidar Amangeldi, Angsar Taigonyrov, Muhammad Huzaid Jawad, Chinedu Emmanuel Mbonu

This study evaluates the trade-offs between convolutional and transformer-based architectures on both medical and general-purpose image classification benchmarks. We use ResNet-18 as our baseline and introduce a fine-tuning strategy applied to four Vision Transformer variants (Tiny, Small, Base, Large) on DermatologyMNIST and TinyImageNet. Our goal is to reduce inference latency and model complexity with acceptable accuracy degradation. Through systematic hyperparameter variations, we demonstrate that appropriately fine-tuned Vision Transformers can match or exceed the baseline's performance, achieve faster inference, and operate with fewer parameters, highlighting their viability for deployment in resource-constrained environments.

📄 PDF Abstract BibTeX arXiv:2505.08259

Code (0)

등록된 구현이 없습니다.

Tasks

image-classificationImage Classification

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Multi-Head Attention 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Vision Transformer The Vision Transformer, or ViT, is a model for image classification that employs a Transformer-like architecture over…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Position-Wise Feed-Forward Layer 설명 없음

Similar Papers 제목 키워드 기반

Improving Diagnostic Accuracy of Pigmented Skin Lesions With CNNs: an Application on the DermaMNIST Dataset

2025-07-17 · Nerma Kadric, Amila Akagic, Medina Kapo arxiv

Pigmented skin lesions represent localized areas of increased melanin and can indicate serious conditions like melanoma, a major contributor to skin cancer mortality. The MedMNIST v2 dataset, inspired by MNIST, was recen…

Multi-class ClassificationTransfer Learning

Vision Transformers in 2022: An Update on Tiny ImageNet

2022-05-21 · Ethan Huynh

The recent advances in image transformers have shown impressive results and have largely closed the gap between traditional CNN architectures. The standard procedure is to train on large datasets like ImageNet-21k and th…

Image ClassificationTransfer Learning

Comparative Analysis of Lightweight Deep Learning Models for Memory-Constrained Devices

2025-05-06 · Tasnim Shahriar

This paper presents a comprehensive evaluation of lightweight deep learning models for image classification, emphasizing their suitability for deployment in resource-constrained environments such as low-memory devices. F…

Computational EfficiencyData AugmentationEdge-computingimage-classification+2

Generative Dataset Distillation Based on Diffusion Model

2024-08-16 · Duo Su, Junjie Hou, Guang Li, Ren Togo 외

This paper presents our method for the generative track of The First Dataset Distillation Challenge at ECCV 2024. Since the diffusion model has become the mainstay of generative models because of its high-quality generat…

Data AugmentationDataset Distillationmodel

Tiny Models are the Computational Saver for Large Models

2024-03-26 · Qingyuan Wang, Barry Cardiff, Antoine Frappé, Benoit Larras 외

This paper introduces TinySaver, an early-exit-like dynamic model compression approach which employs tiny models to substitute large models adaptively. Distinct from traditional compression techniques, dynamic methods li…

Computational EfficiencyImage ClassificationModel Compression