Zero-Shot Transfer Image Classification
16개 벤치마크 · 논문 19편 · 이 태스크의 논문 보기 →
Benchmarks
ImageNet
ImageNet V2
ImageNet-A
ImageNet-R
ObjectNet
ImageNet-Sketch
Food-101
SUN
CN-ImageNet
aYahoo
CN-ImageNet V2
CN-ImageNet-A
CN-ImageNet-R
CN-ImageNet-Sketch
ImageNet ReaL
ImageNet-S
Most implemented
Learning Transferable Visual Models From Natural Language Supervision
CoCa: Contrastive Captioners are Image-Text Foundation Models
LiT: Zero-Shot Transfer with Locked-image text Tuning
Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Your Diffusion Model is Secretly a Zero-Shot Classifier
EVA-CLIP: Improved Training Techniques for CLIP at Scale
Papers
EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters
Scaling up contrastive language-image pretraining (CLIP) is critical for empowering both vision and multimodal models. We present EVA-CLIP-18B, the largest and most powerful open-source CLIP model to date, with 18-billio…
image-classificationImage ClassificationZero-Shot Transfer Image ClassificationM2-Encoder: Advancing Bilingual Image-Text Understanding by Large-scale Efficient Pretraining
Vision-language foundation models like CLIP have revolutionized the field of artificial intelligence. Nevertheless, VLM models supporting multi-language, e.g., in both Chinese and English, have lagged due to the relative…
GPUzero-shot-classificationZero-Shot Cross-Modal RetrievalZero-shot Image Retrieval+3InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
The exponential growth of large language models (LLMs) has opened up numerous possibilities for multimodal AGI systems. However, the progress in vision and vision-language foundation models, which are also critical eleme…
Image RetrievalImage-to-Text RetrievalLanguage ModellingLarge Language Model+11Distilling Large Vision-Language Model with Out-of-Distribution Generalizability
Large vision-language models have achieved outstanding performance, but their size and computational requirements make their deployment on resource-constrained devices and time-sensitive tasks impractical. Model distilla…
Few-Shot Image ClassificationImage ClassificationKnowledge DistillationLanguage Modeling+7Alternating Gradient Descent and Mixture-of-Experts for Integrated Multimodal Perception
We present Integrated Multimodal Perception (IMP), a simple and scalable multimodal multi-task training and modeling approach. IMP integrates multimodal inputs including image, video, text, and audio into a single Transf…
Classificationimage-classificationImage ClassificationMixture-of-Experts+7Your Diffusion Model is Secretly a Zero-Shot Classifier
The recent wave of large-scale text-to-image diffusion models has dramatically increased our text-based image generation abilities. These models can generate realistic images for a staggering variety of prompts and exhib…
Domain GeneralizationFine-Grained Image ClassificationImage ClassificationImage Generation+5