Fine-Grained Visual Recognition
4개 벤치마크 · 논문 86편 · 이 태스크의 논문 보기 →
Benchmarks
Most implemented
Fashionpedia: Ontology, Segmentation, and an Attribute Localization Dataset
Bilinear CNNs for Fine-grained Visual Recognition
Deep CNNs Meet Global Covariance Pooling: Better Representation and Generalization
Retrieving Similar E-Commerce Images Using Deep Learning
RAR: Retrieving And Ranking Augmented MLLMs for Visual Recognition
CLIP-Art: Contrastive Pre-training for Fine-Grained Art Classification
Papers
ConCA: Concentration-Aware Channel Attention for Fine-Grained Visual Recognition
Lightweight channel attention mechanisms are widely used in image classification, yet their effectiveness in fine-grained visual recognition (FGVR) remains limited. Most modules summarize each channel by global average p…
Fine-Grained Visual RecognitionImage ClassificationSubtoken Vision Transformer for Fine-grained Recognition
We present Subtoken Vision Transformer (SubViT), a selective image tokenization method for fine-grained visual recognition. Standard Vision Transformers compress each fixed-size patch into a single token, although fine-g…
Fine-Grained Visual RecognitionFrequency-Enhanced Dual-Subspace Networks for Few-Shot Fine-Grained Image Classification
Few-shot fine-grained image classification aims to recognize subcategories with high visual similarity using only a limited number of annotated samples. Existing metric learning-based methods typically rely solely on spa…
Fine-Grained Image ClassificationFine-Grained Visual RecognitionComputational EfficiencyMetric LearningProgressive Deep Learning for Automated Spheno-Occipital Synchondrosis Maturation Assessment
Accurate assessment of spheno-occipital synchondrosis (SOS) maturation is a key indicator of craniofacial growth and a critical determinant for orthodontic and surgical timing. However, SOS staging from cone-beam CT (CBC…
Fine-Grained Visual RecognitionRepresentation LearningSARE: Sample-wise Adaptive Reasoning for Training-free Fine-grained Visual Recognition
Recent advances in Large Vision-Language Models (LVLMs) have enabled training-free Fine-Grained Visual Recognition (FGVR). However, effectively exploiting LVLMs for FGVR remains challenging due to the inherent visual amb…
Fine-Grained Visual RecognitionTaxonomy-Aware Representation Alignment for Hierarchical Visual Recognition with Large Multimodal Models
A high-performing, general-purpose visual understanding model should map visual inputs to a taxonomic tree of labels, identify novel categories beyond the training set for which few or no publicly available images exist.…
Fine-Grained Visual RecognitionContrastive Learning