Fine-Grained Image Recognition
4개 벤치마크 · 논문 79편 · 이 태스크의 논문 보기 →
Benchmarks
Most implemented
Multi-branch and Multi-scale Attention Learning for Fine-Grained Visual Categorization
Towards Faster Training of Global Covariance Pooling Networks by Iterative Matrix Square Root Normalization
Learning Multi-Attention Convolutional Neural Network for Fine-Grained Image Recognition
Hawkeye: A PyTorch-based Library for Fine-Grained Image Recognition with Deep Learning
PaLI-X: On Scaling up a Multilingual Vision and Language Model
Papers
Structured-Condensed Prompt Tuning in Vision-Language Models for Fine-grained Image Recognition
Fine-grained image recognition poses a significant challenge due to the substantial expertise and effort required for manual annotation. Vision-language models (VLMs) like CLIP provide a compelling zero-shot alternative,…
Fine-Grained Image RecognitionA Large-Scale Study on the Accuracy vs Cost Trade-offs of Training and Evaluation Settings in Fine-Grained Image Recognition
Prior work on fine-grained image recognition (FGIR) has established the importance of the backbone selection, but has neglected the accuracy-vs-cost trade-offs under different training and evaluation settings. In this wo…
Fine-Grained Image RecognitionData AugmentationHow to Choose Your Teacher for Fine Grained Image Recognition
Fine-grained image recognition classifies subcategories such as bird species or car models. While state-of-the-art (SOTA) models are accurate, they are often too resource-intensive for deployment on constrained devices. …
Fine-Grained Image RecognitionKnowledge DistillationRényi Attention Entropy for Patch Pruning
Transformers are strong baselines in both vision and language because self-attention captures long-range dependencies across tokens. However, the cost of self-attention grows quadratically with the number of tokens. Patc…
Fine-Grained Image RecognitionThinking Beyond Labels: Vocabulary-Free Fine-Grained Recognition using Reasoning-Augmented LMMs
Vocabulary-free fine-grained image recognition aims to distinguish visually similar categories within a meta-class without a fixed, human-defined label set. Existing solutions for this problem are limited by either the u…
Fine-Grained Visual RecognitionFine-Grained Image RecognitionDiVE-k: Differential Visual Reasoning for Fine-grained Image Recognition
Large Vision Language Models (LVLMs) possess extensive text knowledge but struggles to utilize this knowledge for fine-grained image recognition, often failing to differentiate between visually similar categories. Existi…
Fine-Grained Image RecognitionReinforcement LearningVisual Reasoning