paper-with-me

Papers Zero-Shot Transfer Image Classification

“Zero-Shot Transfer Image Classification” 태그가 달린 논문 19편 · 필터 해제

EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

2024-02-06 · Quan Sun, Jinsheng Wang, Qiying Yu, Yufeng Cui 외

Scaling up contrastive language-image pretraining (CLIP) is critical for empowering both vision and multimodal models. We present EVA-CLIP-18B, the largest and most powerful open-source CLIP model to date, with 18-billio…

image-classificationImage ClassificationZero-Shot Transfer Image Classification

M2-Encoder: Advancing Bilingual Image-Text Understanding by Large-scale Efficient Pretraining

2024-01-29 · Qingpei Guo, Furong Xu, Hanxiao Zhang, Wang Ren 외

Vision-language foundation models like CLIP have revolutionized the field of artificial intelligence. Nevertheless, VLM models supporting multi-language, e.g., in both Chinese and English, have lagged due to the relative…

GPUzero-shot-classificationZero-Shot Cross-Modal RetrievalZero-shot Image Retrieval+3

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

2023-12-21 · CVPR 2024 1 · Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su 외

The exponential growth of large language models (LLMs) has opened up numerous possibilities for multimodal AGI systems. However, the progress in vision and vision-language foundation models, which are also critical eleme…

Image RetrievalImage-to-Text RetrievalLanguage ModellingLarge Language Model+11

Distilling Large Vision-Language Model with Out-of-Distribution Generalizability

2023-07-06 · ICCV 2023 1 · Xuanlin Li, Yunhao Fang, Minghua Liu, Zhan Ling 외

Large vision-language models have achieved outstanding performance, but their size and computational requirements make their deployment on resource-constrained devices and time-sensitive tasks impractical. Model distilla…

Few-Shot Image ClassificationImage ClassificationKnowledge DistillationLanguage Modeling+7

Alternating Gradient Descent and Mixture-of-Experts for Integrated Multimodal Perception

2023-05-10 · NeurIPS 2023 11 · Hassan Akbari, Dan Kondratyuk, Yin Cui, Rachel Hornung 외

We present Integrated Multimodal Perception (IMP), a simple and scalable multimodal multi-task training and modeling approach. IMP integrates multimodal inputs including image, video, text, and audio into a single Transf…

Classificationimage-classificationImage ClassificationMixture-of-Experts+7

Your Diffusion Model is Secretly a Zero-Shot Classifier

2023-03-28 · ICCV 2023 1 · Alexander C. Li, Mihir Prabhudesai, Shivam Duggal, Ellis Brown 외

The recent wave of large-scale text-to-image diffusion models has dramatically increased our text-based image generation abilities. These models can generate realistic images for a staggering variety of prompts and exhib…

Domain GeneralizationFine-Grained Image ClassificationImage ClassificationImage Generation+5

EVA-CLIP: Improved Training Techniques for CLIP at Scale

2023-03-27 · Quan Sun, Yuxin Fang, Ledell Wu, Xinlong Wang 외

Contrastive language-image pre-training, CLIP for short, has gained increasing attention for its potential in various scenarios. In this paper, we propose EVA-CLIP, a series of models that significantly improve the effic…

Image ClassificationRepresentation LearningZero-Shot Action RecognitionZero-Shot Transfer Image Classification

The effectiveness of MAE pre-pretraining for billion-scale pretraining

2023-03-23 · ICCV 2023 1 · Mannat Singh, Quentin Duval, Kalyan Vasudev Alwala, Haoqi Fan 외

This paper revisits the standard pretrain-then-finetune paradigm used in computer vision for visual recognition tasks. Typically, state-of-the-art foundation models are pretrained using large scale (weakly) supervised da…

Action ClassificationAction RecognitionFew-Shot Image Classificationimage-classification+6

Scaling Vision Transformers to 22 Billion Parameters

2023-02-10 · Mostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski 외

The scaling of Transformers has driven breakthrough capabilities for language models. At present, the largest large language models (LLMs) contain upwards of 100B parameters. Vision Transformers (ViT) have introduced the…

Action ClassificationFairnessImage ClassificationLinear-Probe Classification+2

Learning Customized Visual Models with Retrieval-Augmented Knowledge

2023-01-17 · CVPR 2023 1 · Haotian Liu, Kilho Son, Jianwei Yang, Ce Liu 외

Image-text contrastive learning models such as CLIP have demonstrated strong task transfer ability. The high generality and usability of these visual models is achieved via a web-scale data collection process to ensure b…

Contrastive LearningRetrievalSemi-Supervised Image Classificationzero-shot-classification+2

AltCLIP: Altering the Language Encoder in CLIP for Extended Language Capabilities

2022-11-12 · Zhongzhi Chen, Guang Liu, Bo-Wen Zhang, Fulong Ye 외

In this work, we present a conceptually simple and effective method to train a strong bilingual/multilingual multimodal representation model. Starting from the pre-trained multimodal representation model CLIP released by…

Contrastive LearningCross-Modal RetrievalImage ClassificationImage Retrieval+9

PaLI: A Jointly-Scaled Multilingual Language-Image Model

2022-09-14 · Xi Chen, Xiao Wang, Soravit Changpinyo, AJ Piergiovanni 외

Effective scaling and a flexible task interface enable large language models to excel at many tasks. We present PaLI (Pathways Language and Image model), a model that extends this approach to the joint modeling of langua…

DecoderFew-Shot Image ClassificationImage CaptioningImage Classification+7

CoCa: Contrastive Captioners are Image-Text Foundation Models

2022-05-04 · Jiahui Yu, ZiRui Wang, Vijay Vasudevan, Legg Yeung 외

Exploring large-scale pretrained foundation models is of significant interest in computer vision because these models can be quickly transferred to many downstream tasks. This paper presents Contrastive Captioner (CoCa),…

Action ClassificationDecoderImage CaptioningImage Classification+9

Florence: A New Foundation Model for Computer Vision

2021-11-22 · Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella 외

Automated visual understanding of our diverse and open world demands computer vision models to generalize well with minimal customization for specific tasks, similar to human vision. Computer vision foundation models, wh…

Action ClassificationAction RecognitionAction Recognition In VideosCross-Modal Retrieval+15

Combined Scaling for Zero-shot Transfer Learning

2021-11-19 · Hieu Pham, Zihang Dai, Golnaz Ghiasi, Kenji Kawaguchi 외

We present a combined scaling method - named BASIC - that achieves 85.7% top-1 accuracy on the ImageNet ILSVRC-2012 validation set without learning from any labeled ImageNet example. This accuracy surpasses best publishe…

ClassificationContrastive LearningImage ClassificationTransfer Learning+1

LiT: Zero-Shot Transfer with Locked-image text Tuning

2021-11-15 · CVPR 2022 1 · Xiaohua Zhai, Xiao Wang, Basil Mustafa, Andreas Steiner 외

This paper presents contrastive-tuning, a simple method employing contrastive training to align image and text models while still taking advantage of their pre-training. In our empirical study we find that locked pre-tra…

image-classificationImage ClassificationRetrievalZero-Shot Image Classification+1

Learning Transferable Visual Models From Natural Language Supervision

2021-02-26 · Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 외

State-of-the-art computer vision systems are trained to predict a fixed set of predetermined object categories. This restricted form of supervision limits their generality and usability since additional labeled data is n…

Action RecognitionBenchmarkingFew-Shot Image Classificationgeo-localization+21

Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

2021-02-11 · Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 외

Pre-trained representations are becoming crucial for many NLP and perception tasks. While representation learning in NLP has transitioned to training on raw text without human annotations, visual and vision-language repr…

Cross-Modal RetrievalFine-Grained Image Classificationimage-classificationImage Classification+7

Learning Visual N-Grams from Web Data

2016-12-29 · ICCV 2017 10 · Ang Li, Allan Jabri, Armand Joulin, Laurens van der Maaten

Real-world image recognition systems need to recognize tens of thousands of classes that constitute a plethora of visual concepts. The traditional approach of annotating thousands of images per class for training is infe…

Language ModelingLanguage ModellingRepresentation LearningRetrieval+1
1–19 / 19