Papers Zero-shot Text-to-Image Retrieval
“Zero-shot Text-to-Image Retrieval” 태그가 달린 논문 15편 · 필터 해제
An analysis of vision-language models for fabric retrieval
Effective cross-modal retrieval is essential for applications like information retrieval and recommendation systems, particularly in specialized domains such as manufacturing, where product information often consists of …
AttributeCross-Modal RetrievalImage RetrievalInformation Retrieval+3CLIP-PING: Boosting Lightweight Vision-Language Models with Proximus Intrinsic Neighbors Guidance
Beyond the success of Contrastive Language-Image Pre-training (CLIP), recent trends mark a shift toward exploring the applicability of lightweight vision-language models for resource-constrained scenarios. These models o…
Contrastive Learningcross-modal alignmentCross-Modal RetrievalLinear evaluation+6MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions
Image retrieval, i.e., finding desired images given a reference image, inherently encompasses rich, multi-faceted search intents that are difficult to capture solely using image-based measures. Recent works leverage text…
Image RetrievalImplicit RelationsRetrievalSupervised Image Retrieval+2M2-Encoder: Advancing Bilingual Image-Text Understanding by Large-scale Efficient Pretraining
Vision-language foundation models like CLIP have revolutionized the field of artificial intelligence. Nevertheless, VLM models supporting multi-language, e.g., in both Chinese and English, have lagged due to the relative…
GPUzero-shot-classificationZero-Shot Cross-Modal RetrievalZero-shot Image Retrieval+3Linguistic-Aware Patch Slimming Framework for Fine-grained Cross-Modal Alignment
Cross-modal alignment aims to build a bridge connecting vision and language. It is an important multi-modal task that efficiently learns the semantic similarities between images and texts. Traditional fine-grained al…
cross-modal alignmentCross-Modal RetrievalImage RetrievalImage-to-Text Retrieval+4ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities
In this work, we explore a scalable way for building a general representation model toward unlimited modalities. We release ONE-PEACE, a highly extensible model with 4B parameters that can seamlessly align and integrate …
1 Image, 2*2 StitchiAction ClassificationAudioCapsAudio Classification+18CAVL: Learning Contrastive and Adaptive Representations of Vision and Language
Visual and linguistic pre-training aims to learn vision and language representations together, which can be transferred to visual-linguistic downstream tasks. However, there exists semantic confusion between language and…
Image RetrievalPhrase GroundingQuestion AnsweringRetrieval+6Sigmoid Loss for Language Image Pre-Training
We propose a simple pairwise Sigmoid loss for Language-Image Pre-training (SigLIP). Unlike standard contrastive learning with softmax normalization, the sigmoid loss operates solely on image-text pairs and does not requi…
Contrastive LearningDisentanglementImage-to-Text RetrievalZero-shot Text-to-Image RetrievalBLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
The cost of vision-and-language pre-training has become increasingly prohibitive due to end-to-end training of large-scale models. This paper proposes BLIP-2, a generic and efficient pre-training strategy that bootstraps…
Generative Visual Question AnsweringImage CaptioningImage RetrievalImage to text+13Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese
The tremendous success of CLIP (Radford et al., 2021) has promoted the research and application of contrastive learning for vision-language pretraining. In this work, we construct a large-scale dataset of image-text pair…
Contrastive Learningimage-classificationImage ClassificationImage Retrieval+7ERNIE-ViL 2.0: Multi-view Contrastive Learning for Image-Text Pre-training
Recent Vision-Language Pre-trained (VLP) models based on dual encoder have attracted extensive attention from academia and industry due to their superior performance on various cross-modal tasks and high computational ef…
Computational EfficiencyContrastive LearningCross-Modal RetrievalImage Retrieval+5Crossmodal-3600: A Massively Multilingual Multimodal Evaluation Dataset
Research in massively multilingual image captioning has been severely hampered by a lack of high-quality evaluation datasets. In this paper we present the Crossmodal-3600 dataset (XM3600 in short), a geographically diver…
Image CaptioningImage RetrievalImage-text RetrievalImage-to-Text Retrieval+3FLAVA: A Foundational Language And Vision Alignment Model
State-of-the-art vision and vision-and-language models rely on large-scale visio-linguistic pretraining for obtaining good performance on a variety of downstream tasks. Generally, such models are often either cross-modal…
Image RetrievalImage-to-Text RetrievalVisual ReasoningZero-shot Image Retrieval+2Learning Transferable Visual Models From Natural Language Supervision
State-of-the-art computer vision systems are trained to predict a fixed set of predetermined object categories. This restricted form of supervision limits their generality and usability since additional labeled data is n…
Action RecognitionBenchmarkingFew-Shot Image Classificationgeo-localization+21ZSCRGAN: A GAN-based Expectation Maximization Model for Zero-Shot Retrieval of Images from Textual Descriptions
Most existing algorithms for cross-modal Information Retrieval are based on a supervised train-test setup, where a model learns to align the mode of the query (e.g., text) to the mode of the documents (e.g., images) from…
Cross-Modal Information RetrievalImage RetrievalInformation RetrievalRetrieval+3