Papers Zero-Shot Transfer Image Classification
“Zero-Shot Transfer Image Classification” 태그가 달린 논문 19편 · 필터 해제
EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters
Scaling up contrastive language-image pretraining (CLIP) is critical for empowering both vision and multimodal models. We present EVA-CLIP-18B, the largest and most powerful open-source CLIP model to date, with 18-billio…
image-classificationImage ClassificationZero-Shot Transfer Image ClassificationM2-Encoder: Advancing Bilingual Image-Text Understanding by Large-scale Efficient Pretraining
Vision-language foundation models like CLIP have revolutionized the field of artificial intelligence. Nevertheless, VLM models supporting multi-language, e.g., in both Chinese and English, have lagged due to the relative…
GPUzero-shot-classificationZero-Shot Cross-Modal RetrievalZero-shot Image Retrieval+3InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
The exponential growth of large language models (LLMs) has opened up numerous possibilities for multimodal AGI systems. However, the progress in vision and vision-language foundation models, which are also critical eleme…
Image RetrievalImage-to-Text RetrievalLanguage ModellingLarge Language Model+11Distilling Large Vision-Language Model with Out-of-Distribution Generalizability
Large vision-language models have achieved outstanding performance, but their size and computational requirements make their deployment on resource-constrained devices and time-sensitive tasks impractical. Model distilla…
Few-Shot Image ClassificationImage ClassificationKnowledge DistillationLanguage Modeling+7Alternating Gradient Descent and Mixture-of-Experts for Integrated Multimodal Perception
We present Integrated Multimodal Perception (IMP), a simple and scalable multimodal multi-task training and modeling approach. IMP integrates multimodal inputs including image, video, text, and audio into a single Transf…
Classificationimage-classificationImage ClassificationMixture-of-Experts+7Your Diffusion Model is Secretly a Zero-Shot Classifier
The recent wave of large-scale text-to-image diffusion models has dramatically increased our text-based image generation abilities. These models can generate realistic images for a staggering variety of prompts and exhib…
Domain GeneralizationFine-Grained Image ClassificationImage ClassificationImage Generation+5EVA-CLIP: Improved Training Techniques for CLIP at Scale
Contrastive language-image pre-training, CLIP for short, has gained increasing attention for its potential in various scenarios. In this paper, we propose EVA-CLIP, a series of models that significantly improve the effic…
Image ClassificationRepresentation LearningZero-Shot Action RecognitionZero-Shot Transfer Image ClassificationThe effectiveness of MAE pre-pretraining for billion-scale pretraining
This paper revisits the standard pretrain-then-finetune paradigm used in computer vision for visual recognition tasks. Typically, state-of-the-art foundation models are pretrained using large scale (weakly) supervised da…
Action ClassificationAction RecognitionFew-Shot Image Classificationimage-classification+6Scaling Vision Transformers to 22 Billion Parameters
The scaling of Transformers has driven breakthrough capabilities for language models. At present, the largest large language models (LLMs) contain upwards of 100B parameters. Vision Transformers (ViT) have introduced the…
Action ClassificationFairnessImage ClassificationLinear-Probe Classification+2Learning Customized Visual Models with Retrieval-Augmented Knowledge
Image-text contrastive learning models such as CLIP have demonstrated strong task transfer ability. The high generality and usability of these visual models is achieved via a web-scale data collection process to ensure b…
Contrastive LearningRetrievalSemi-Supervised Image Classificationzero-shot-classification+2AltCLIP: Altering the Language Encoder in CLIP for Extended Language Capabilities
In this work, we present a conceptually simple and effective method to train a strong bilingual/multilingual multimodal representation model. Starting from the pre-trained multimodal representation model CLIP released by…
Contrastive LearningCross-Modal RetrievalImage ClassificationImage Retrieval+9PaLI: A Jointly-Scaled Multilingual Language-Image Model
Effective scaling and a flexible task interface enable large language models to excel at many tasks. We present PaLI (Pathways Language and Image model), a model that extends this approach to the joint modeling of langua…
DecoderFew-Shot Image ClassificationImage CaptioningImage Classification+7CoCa: Contrastive Captioners are Image-Text Foundation Models
Exploring large-scale pretrained foundation models is of significant interest in computer vision because these models can be quickly transferred to many downstream tasks. This paper presents Contrastive Captioner (CoCa),…
Action ClassificationDecoderImage CaptioningImage Classification+9Florence: A New Foundation Model for Computer Vision
Automated visual understanding of our diverse and open world demands computer vision models to generalize well with minimal customization for specific tasks, similar to human vision. Computer vision foundation models, wh…
Action ClassificationAction RecognitionAction Recognition In VideosCross-Modal Retrieval+15Combined Scaling for Zero-shot Transfer Learning
We present a combined scaling method - named BASIC - that achieves 85.7% top-1 accuracy on the ImageNet ILSVRC-2012 validation set without learning from any labeled ImageNet example. This accuracy surpasses best publishe…
ClassificationContrastive LearningImage ClassificationTransfer Learning+1LiT: Zero-Shot Transfer with Locked-image text Tuning
This paper presents contrastive-tuning, a simple method employing contrastive training to align image and text models while still taking advantage of their pre-training. In our empirical study we find that locked pre-tra…
image-classificationImage ClassificationRetrievalZero-Shot Image Classification+1Learning Transferable Visual Models From Natural Language Supervision
State-of-the-art computer vision systems are trained to predict a fixed set of predetermined object categories. This restricted form of supervision limits their generality and usability since additional labeled data is n…
Action RecognitionBenchmarkingFew-Shot Image Classificationgeo-localization+21Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Pre-trained representations are becoming crucial for many NLP and perception tasks. While representation learning in NLP has transitioned to training on raw text without human annotations, visual and vision-language repr…
Cross-Modal RetrievalFine-Grained Image Classificationimage-classificationImage Classification+7Learning Visual N-Grams from Web Data
Real-world image recognition systems need to recognize tens of thousands of classes that constitute a plethora of visual concepts. The traditional approach of annotating thousands of images per class for training is infe…
Language ModelingLanguage ModellingRepresentation LearningRetrieval+1