Papers Zero-shot Image Retrieval
“Zero-shot Image Retrieval” 태그가 달린 논문 29편 · 필터 해제
Revisiting CLIP: Efficient Alignment of 3D MRI and Tabular Data using Domain-Specific Foundation Models
Multi-modal models require aligned, shared embedding spaces. However, common CLIP-based approaches need large amounts of samples and do not natively support 3D or tabular data, both of which are crucial in the medical do…
Image RetrievalRetrievalzero-shot-classificationZero-shot Image Retrieval+1CLIP-PING: Boosting Lightweight Vision-Language Models with Proximus Intrinsic Neighbors Guidance
Beyond the success of Contrastive Language-Image Pre-training (CLIP), recent trends mark a shift toward exploring the applicability of lightweight vision-language models for resource-constrained scenarios. These models o…
Contrastive Learningcross-modal alignmentCross-Modal RetrievalLinear evaluation+6Piecewise-Linear Manifolds for Deep Metric Learning
Unsupervised deep metric learning (UDML) focuses on learning a semantic representation space using only unlabeled data. This challenging problem requires accurately estimating the similarity between data points, which is…
Image RetrievalMetric LearningRetrievalZero-shot Image RetrievalM2-Encoder: Advancing Bilingual Image-Text Understanding by Large-scale Efficient Pretraining
Vision-language foundation models like CLIP have revolutionized the field of artificial intelligence. Nevertheless, VLM models supporting multi-language, e.g., in both Chinese and English, have lagged due to the relative…
GPUzero-shot-classificationZero-Shot Cross-Modal RetrievalZero-shot Image Retrieval+3InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
The exponential growth of large language models (LLMs) has opened up numerous possibilities for multimodal AGI systems. However, the progress in vision and vision-language foundation models, which are also critical eleme…
Image RetrievalImage-to-Text RetrievalLanguage ModellingLarge Language Model+11Context-I2W: Mapping Images to Context-dependent Words for Accurate Zero-Shot Composed Image Retrieval
Different from Composed Image Retrieval task that requires expensive labels for training task-specific models, Zero-Shot Composed Image Retrieval (ZS-CIR) involves diverse tasks with a broad range of visual content manip…
AttributeImage RetrievalObjectRetrieval+2GrowCLIP: Data-aware Automatic Model Growing for Large-scale Contrastive Language-Image Pre-training
Cross-modal pre-training has shown impressive performance on a wide range of downstream tasks, benefiting from massive image-text pairs collected from the Internet. In practice, online data are growing constantly, highli…
image-classificationImage ClassificationImage RetrievalImage to text+2FACTUAL: A Benchmark for Faithful and Consistent Textual Scene Graph Parsing
Textual scene graph parsing has become increasingly important in various vision-language applications, including image caption evaluation and image retrieval. However, existing scene graph parsers that convert image capt…
Graph SimilarityHuman Judgment CorrelationImage CaptioningImage Retrieval+2Pic2Word: Mapping Pictures to Words for Zero-shot Composed Image Retrieval
In Composed Image Retrieval (CIR), a user combines a query image with text to describe their intended target. Existing methods rely on supervised learning of CIR models using labeled triplets consisting of the query imag…
AttributeComposed Image Retrieval (CoIR)Image RetrievalRetrieval+2AltCLIP: Altering the Language Encoder in CLIP for Extended Language Capabilities
In this work, we present a conceptually simple and effective method to train a strong bilingual/multilingual multimodal representation model. Starting from the pre-trained multimodal representation model CLIP released by…
Contrastive LearningCross-Modal RetrievalImage ClassificationImage Retrieval+9Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese
The tremendous success of CLIP (Radford et al., 2021) has promoted the research and application of contrastive learning for vision-language pretraining. In this work, we construct a large-scale dataset of image-text pair…
Contrastive Learningimage-classificationImage ClassificationImage Retrieval+7General Image Descriptors for Open World Image Retrieval using ViT CLIP
The Google Universal Image Embedding (GUIE) Challenge is one of the first competitions in multi-domain image representations in the wild, covering a wide distribution of objects: landmarks, artwork, food, etc. This is a …
Image RetrievalRetrievalZero-Shot Image ClassificationZero-shot Image Retrieval+1ERNIE-ViL 2.0: Multi-view Contrastive Learning for Image-Text Pre-training
Recent Vision-Language Pre-trained (VLP) models based on dual encoder have attracted extensive attention from academia and industry due to their superior performance on various cross-modal tasks and high computational ef…
Computational EfficiencyContrastive LearningCross-Modal RetrievalImage Retrieval+5FETA: Towards Specializing Foundation Models for Expert Task Applications
Foundation Models (FMs) have demonstrated unprecedented capabilities including zero-shot learning, high fidelity data synthesis, and out of domain generalization. However, as we show in this paper, FMs still have poor ou…
Domain GeneralizationFew-Shot LearningImage RetrievalImage-text Retrieval+7Curriculum Learning for Data-Efficient Vision-Language Alignment
Aligning image and text encoders from scratch using contrastive learning requires large amounts of paired image-text data. We alleviate this need by aligning individually pre-trained language and vision representation mo…
Contrastive LearningImage RetrievalObjectRetrieval+1Cross-lingual and Multilingual CLIP
The long-standing endeavor of relating the textual and the visual domain recently underwent a pivotal breakthrough, as OpenAI released CLIP. This model distinguishes how well an English text corresponds with a given imag…
Contrastive LearningImage-text RetrievalMachine TranslationRetrieval+2CCMB: A Large-scale Chinese Cross-modal Benchmark
Vision-language pre-training (VLP) on large-scale datasets has shown premier performance on various downstream tasks. In contrast to plenty of available benchmarks with English corpus, large-scale pre-training datasets a…
image-classificationImage ClassificationImage GenerationImage Retrieval+9Wukong: A 100 Million Large-scale Chinese Cross-modal Pre-training Benchmark
Vision-Language Pre-training (VLP) models have shown remarkable performance on various downstream tasks. Their success heavily relies on the scale of pre-trained cross-modal datasets. However, the lack of large-scale dat…
BenchmarkingContrastive Learningimage-classificationImage Classification+6Visual Representation Learning with Self-Supervised Attention for Low-Label High-data Regime
Self-supervision has shown outstanding results for natural language processing, and more recently, for image recognition. Simultaneously, vision transformers and its variants have emerged as a promising and scalable alte…
Few-Shot Image Classificationimage-classificationImage ClassificationImage Retrieval+4FLAVA: A Foundational Language And Vision Alignment Model
State-of-the-art vision and vision-and-language models rely on large-scale visio-linguistic pretraining for obtaining good performance on a variety of downstream tasks. Generally, such models are often either cross-modal…
Image RetrievalImage-to-Text RetrievalVisual ReasoningZero-shot Image Retrieval+2