arXiv · 2505.08259
CNN and ViT Efficiency Study on Tiny ImageNet and DermaMNIST Datasets
Abstract
This study evaluates the trade-offs between convolutional and transformer-based architectures on both medical and general-purpose image classification benchmarks. We use ResNet-18 as our baseline and introduce a fine-tuning strategy applied to four Vision Transformer variants (Tiny, Small, Base, Large) on DermatologyMNIST and TinyImageNet. Our goal is to reduce inference latency and model complexity with acceptable accuracy degradation. Through systematic hyperparameter variations, we demonstrate that appropriately fine-tuned Vision Transformers can match or exceed the baseline's performance, achieve faster inference, and operate with fewer parameters, highlighting their viability for deployment in resource-constrained environments.
Explore related subjects
Keep this discovery
Aidar Amangeldi, Angsar Taigonyrov, Muhammad Huzaifa Jawad, Chinedu Emmanuel Mbonu. 2025-05-13. CNN and ViT Efficiency Study on Tiny ImageNet and DermaMNIST Datasets. https://arxiv.org/abs/2505.08259
Cite the original work for its findings. Save a collection to share your selection of sources.