Papers Few-Shot Image Classification
“Few-Shot Image Classification” 태그가 달린 논문 367편 · 필터 해제
Decompose, Compare, and Decide: Multimodal LLMs are Implicit Few-Shot Learners
Multimodal Large Language Models (MLLMs) have demonstrated remarkable abilities when analyzing images, yet translating these capabilities to few-shot image classification remains challenging. To bridge this gap, we prese…
Few-Shot Image ClassificationHippocampus-DETR: An Explicit Memory Object Detection Framework Based on Hippocampus Modeling
This paper addresses the lack of explicit memory mechanisms in current object detection models and proposes Hippocampus-DETR, a novel detection framework based on biological hippocampal memory modeling. This framework in…
Few-Shot Image ClassificationNovel Object DetectionImage RestorationMAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models
Adapting large vision-language models (VLMs) such as CLIP to downstream tasks remains challenging, as full fine-tuning is computationally prohibitive and prone to overfitting in low-data regimes. Parameter-efficient fine…
parameter-efficient fine-tuningFew-Shot Image ClassificationComputational EfficiencyA$_3$B$_2$: Adaptive Asymmetric Adapter for Alleviating Branch Bias in Vision-Language Image Classification with Few-Shot Learning
Efficient transfer learning methods for large-scale vision-language models ($e.g.$, CLIP) enable strong few-shot transfer, yet existing adaptation methods follow a fixed fine-tuning paradigm that implicitly assumes a uni…
Few-Shot Image ClassificationFew-Shot LearningTransfer LearningSpurAudio: A Benchmark for Studying Shortcut Learning in Few-Shot Audio Classification
Few-shot classification (FSC) is widely used for learning from limited labeled data, yet most evaluations implicitly assume that target concepts are independent of contextual cues. In real-world settings, however, exampl…
Few-Shot Image ClassificationFew-Shot Audio ClassificationCross-Modal Prototype Alignment and Mixing for Training-Free Few-Shot Classification
Vision-language models (VLMs) like CLIP are trained with the objective of aligning text and image pairs. To improve CLIP-based few-shot image classification, recent works have observed that, along with text embeddings, i…
Few-Shot Image ClassificationSemi-Supervised Few-Shot Adaptation of Vision-Language Models
Vision-language models (VLMs) pre-trained on large, heterogeneous data sources are becoming increasingly popular, providing rich multi-modal embeddings that enable efficient transfer to new tasks. A particularly relevant…
Few-Shot Image ClassificationAdapting Multimodal Foundation Models for Few-Shot Learning: A Comprehensive Study on Contrastive Captioners
Large-scale multimodal foundation models, particularly Contrastive Captioners (CoCa), have achieved state-of-the-art results by unifying contrastive alignment with generative captioning. While zero-shot transfer capabili…
parameter-efficient fine-tuningFew-Shot Image ClassificationFew-Shot LearningData AugmentationAdvancing Cache-Based Few-Shot Classification via Patch-Driven Relational Gated Graph Attention
Few-shot image classification remains difficult under limited supervision and visual domain shift. Recent cache-based adaptation approaches (e.g., Tip-Adapter) address this challenge to some extent by learning lightweigh…
Few-Shot Image ClassificationTraining-Free Synthetic Data Generation with Dual IP-Adapter Guidance
Few-shot image classification remains challenging due to the limited availability of labeled examples. Recent approaches have explored generating synthetic training data using text-to-image diffusion models, but often re…
Few-Shot Image ClassificationImage-to-Image TranslationSynthetic Data GenerationDual-View Alignment Learning with Hierarchical-Prompt for Class-Imbalance Multi-Label Classification
Real-world datasets often exhibit class imbalance across multiple categories, manifesting as long-tailed distributions and few-shot scenarios. This is especially challenging in Class-Imbalanced Multi-Label Image Classifi…
Multi-Label Image ClassificationFew-Shot Image ClassificationMulti-Label ClassificationObject RecognitionPreserve and Sculpt: Manifold-Aligned Fine-tuning of Vision-Language Models for Few-Shot Learning
Pretrained vision-language models (VLMs), such as CLIP, have shown remarkable potential in few-shot image classification and led to numerous effective transfer learning strategies. These methods leverage the pretrained k…
Few-Shot Image ClassificationFew-Shot LearningTransfer LearningDomain AdaptationObject-Centric Cropping for Visual Few-Shot Classification
In the domain of Few-Shot Image Classification, operating with as little as one example per class, the presence of image ambiguities stemming from multiple objects or complex backgrounds can significantly deteriorate per…
Few-Shot Image ClassificationProtoConNet: Prototypical Augmentation and Alignment for Open-Set Few-Shot Image Classification
Open-set few-shot image classification aims to train models using a small amount of labeled data, enabling them to achieve good generalization when confronted with unknown environments. Existing methods mainly use visual…
Few-Shot Image ClassificationRepresentation LearningViT-ProtoNet for Few-Shot Image Classification: A Multi-Benchmark Evaluation
The remarkable representational power of Vision Transformers (ViTs) remains underutilized in few-shot image classification. In this work, we introduce ViT-ProtoNet, which integrates a ViT-Small backbone into the Prototyp…
Few-Shot Image Classificationimage-classificationImage ClassificationInterpretable Few-Shot Image Classification via Prototypical Concept-Guided Mixture of LoRA Experts
Self-Explainable Models (SEMs) rely on Prototypical Concept Learning (PCL) to enable their visual recognition processes more interpretable, but they often struggle in data-scarce settings where insufficient training samp…
Explainable ModelsFew-Shot Image Classificationimage-classificationImage ClassificationProvably Improving Generalization of Few-Shot Models with Synthetic Data
Few-shot image classification remains challenging due to the scarcity of labeled training examples. Augmenting them with synthetic data has emerged as a promising way to alleviate this issue, but models trained on synthe…
Few-Shot Image Classificationimage-classificationImage ClassificationSimple Semi-supervised Knowledge Distillation from Vision-Language Models via $\mathbf{\texttt{D}}$ual-$\mathbf{\texttt{H}}$ead $\mathbf{\texttt{O}}$ptimization
Vision-language models (VLMs) have achieved remarkable success across diverse tasks by leveraging rich textual information with minimal labeled data. However, deploying such large models remains challenging, particularly…
Few-Shot Image ClassificationKnowledge DistillationSemi-Supervised Image ClassificationSemi-Supervised Image Classification on ImageNet - 10% labeled dataBrain Inspired Adaptive Memory Dual-Net for Few-Shot Image Classification
Few-shot image classification has become a popular research topic for its wide application in real-world scenarios, however the problem of supervision collapse induced by single image-level annotation remains a major cha…
Few-Shot Image ClassificationHippocampusimage-classificationImage ClassificationInPK: Infusing Prior Knowledge into Prompt for Vision-Language Models
Prompt tuning has become a popular strategy for adapting Vision-Language Models (VLMs) to zero/few-shot visual recognition tasks. Some prompting techniques introduce prior knowledge due to its richness, but when learnabl…
Few-Shot Image Classificationimage-classificationImage Classification