Few-Shot Image Classification
89개 벤치마크 · 논문 367편 · 이 태스크의 논문 보기 →
Benchmarks
Mini-Imagenet 5-way (1-shot)
Mini-Imagenet 5-way (5-shot)
Tiered ImageNet 5-way (5-shot)
CIFAR-FS 5-way (5-shot)
CIFAR-FS 5-way (1-shot)
CUB 200 5-way 1-shot
CUB 200 5-way 5-shot
FC100 5-way (1-shot)
FC100 5-way (5-shot)
Meta-Dataset
OMNIGLOT - 1-Shot, 20-way
OMNIGLOT - 5-Shot, 20-way
OMNIGLOT - 1-Shot, 5-way
OMNIGLOT - 5-Shot, 5-way
Meta-Dataset Rank
Bongard-HOI
ImageNet - 1-shot
ImageNet - 5-shot
ImageNet-FS (2-shot, novel)
ImageNet-FS (5-shot, all)
ImageNet - 10-shot
ImageNet-FS (1-shot, novel)
Stanford Cars 5-way (1-shot)
Stanford Cars 5-way (5-shot)
Stanford Dogs 5-way (5-shot)
CUB-200-2011 - 0-Shot
ImageNet - 0-Shot
CUB 200 50-way (0-shot)
CIFAR100 5-way (1-shot)
ImageNet (1-shot)
SUN - 0-Shot
AWA - 0-Shot
AWA1 - 0-Shot
AWA2 - 0-Shot
CUB 200 5-way
Caltech101
FC100 5-way (10-shot)
Flowers-102 - 0-Shot
Oxford 102 Flower
UT Zappos50K
aPY - 0-Shot
mini-ImageNet - 100-Way
Most implemented
Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks
Learning Transferable Visual Models From Natural Language Supervision
Prototypical Networks for Few-shot Learning
Matching Networks for One Shot Learning
Meta-Dataset: A Dataset of Datasets for Learning to Learn from Few Examples
On First-Order Meta-Learning Algorithms
Papers
Decompose, Compare, and Decide: Multimodal LLMs are Implicit Few-Shot Learners
Multimodal Large Language Models (MLLMs) have demonstrated remarkable abilities when analyzing images, yet translating these capabilities to few-shot image classification remains challenging. To bridge this gap, we prese…
Few-Shot Image ClassificationHippocampus-DETR: An Explicit Memory Object Detection Framework Based on Hippocampus Modeling
This paper addresses the lack of explicit memory mechanisms in current object detection models and proposes Hippocampus-DETR, a novel detection framework based on biological hippocampal memory modeling. This framework in…
Few-Shot Image ClassificationNovel Object DetectionImage RestorationMAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models
Adapting large vision-language models (VLMs) such as CLIP to downstream tasks remains challenging, as full fine-tuning is computationally prohibitive and prone to overfitting in low-data regimes. Parameter-efficient fine…
parameter-efficient fine-tuningFew-Shot Image ClassificationComputational EfficiencyA$_3$B$_2$: Adaptive Asymmetric Adapter for Alleviating Branch Bias in Vision-Language Image Classification with Few-Shot Learning
Efficient transfer learning methods for large-scale vision-language models ($e.g.$, CLIP) enable strong few-shot transfer, yet existing adaptation methods follow a fixed fine-tuning paradigm that implicitly assumes a uni…
Few-Shot Image ClassificationFew-Shot LearningTransfer LearningSpurAudio: A Benchmark for Studying Shortcut Learning in Few-Shot Audio Classification
Few-shot classification (FSC) is widely used for learning from limited labeled data, yet most evaluations implicitly assume that target concepts are independent of contextual cues. In real-world settings, however, exampl…
Few-Shot Image ClassificationFew-Shot Audio ClassificationCross-Modal Prototype Alignment and Mixing for Training-Free Few-Shot Classification
Vision-language models (VLMs) like CLIP are trained with the objective of aligning text and image pairs. To improve CLIP-based few-shot image classification, recent works have observed that, along with text embeddings, i…
Few-Shot Image Classification