paper-with-me

Zero-Shot Transfer Image Classification

16개 벤치마크 · 논문 19편 · 이 태스크의 논문 보기 →

Benchmarks

ImageNet

결과 23개

ImageNet V2

결과 13개

ImageNet-A

결과 13개

ImageNet-R

결과 12개

ObjectNet

결과 9개

ImageNet-Sketch

결과 7개

Food-101

결과 5개

SUN

결과 3개

CN-ImageNet

결과 2개

aYahoo

결과 2개

CN-ImageNet V2

결과 1개

CN-ImageNet-A

결과 1개

CN-ImageNet-R

결과 1개

CN-ImageNet-Sketch

결과 1개

ImageNet ReaL

결과 1개

ImageNet-S

결과 1개

Most implemented

Papers

EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

2024-02-06 · Quan Sun, Jinsheng Wang, Qiying Yu, Yufeng Cui 외

Scaling up contrastive language-image pretraining (CLIP) is critical for empowering both vision and multimodal models. We present EVA-CLIP-18B, the largest and most powerful open-source CLIP model to date, with 18-billio…

image-classificationImage ClassificationZero-Shot Transfer Image Classification

M2-Encoder: Advancing Bilingual Image-Text Understanding by Large-scale Efficient Pretraining

2024-01-29 · Qingpei Guo, Furong Xu, Hanxiao Zhang, Wang Ren 외

Vision-language foundation models like CLIP have revolutionized the field of artificial intelligence. Nevertheless, VLM models supporting multi-language, e.g., in both Chinese and English, have lagged due to the relative…

GPUzero-shot-classificationZero-Shot Cross-Modal RetrievalZero-shot Image Retrieval+3

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

2023-12-21 · CVPR 2024 1 · Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su 외

The exponential growth of large language models (LLMs) has opened up numerous possibilities for multimodal AGI systems. However, the progress in vision and vision-language foundation models, which are also critical eleme…

Image RetrievalImage-to-Text RetrievalLanguage ModellingLarge Language Model+11

Distilling Large Vision-Language Model with Out-of-Distribution Generalizability

2023-07-06 · ICCV 2023 1 · Xuanlin Li, Yunhao Fang, Minghua Liu, Zhan Ling 외

Large vision-language models have achieved outstanding performance, but their size and computational requirements make their deployment on resource-constrained devices and time-sensitive tasks impractical. Model distilla…

Few-Shot Image ClassificationImage ClassificationKnowledge DistillationLanguage Modeling+7

Alternating Gradient Descent and Mixture-of-Experts for Integrated Multimodal Perception

2023-05-10 · NeurIPS 2023 11 · Hassan Akbari, Dan Kondratyuk, Yin Cui, Rachel Hornung 외

We present Integrated Multimodal Perception (IMP), a simple and scalable multimodal multi-task training and modeling approach. IMP integrates multimodal inputs including image, video, text, and audio into a single Transf…

Classificationimage-classificationImage ClassificationMixture-of-Experts+7

Your Diffusion Model is Secretly a Zero-Shot Classifier

2023-03-28 · ICCV 2023 1 · Alexander C. Li, Mihir Prabhudesai, Shivam Duggal, Ellis Brown 외

The recent wave of large-scale text-to-image diffusion models has dramatically increased our text-based image generation abilities. These models can generate realistic images for a staggering variety of prompts and exhib…

Domain GeneralizationFine-Grained Image ClassificationImage ClassificationImage Generation+5

전체 19편 보기 →