paper-with-me

Papers Zero-shot Text-to-Image Retrieval

“Zero-shot Text-to-Image Retrieval” 태그가 달린 논문 15편 · 필터 해제

An analysis of vision-language models for fabric retrieval

2025-07-07 · Francesco Giuliari, Asif Khan Pattan, Mohamed Lamine Mekhalfi, Fabio Poiesi

Effective cross-modal retrieval is essential for applications like information retrieval and recommendation systems, particularly in specialized domains such as manufacturing, where product information often consists of …

AttributeCross-Modal RetrievalImage RetrievalInformation Retrieval+3

CLIP-PING: Boosting Lightweight Vision-Language Models with Proximus Intrinsic Neighbors Guidance

2024-12-05 · Chu Myaet Thwal, Ye Lin Tun, Minh N. H. Nguyen, Eui-Nam Huh 외

Beyond the success of Contrastive Language-Image Pre-training (CLIP), recent trends mark a shift toward exploring the applicability of lightweight vision-language models for resource-constrained scenarios. These models o…

Contrastive Learningcross-modal alignmentCross-Modal RetrievalLinear evaluation+6

MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions

2024-03-28 · Kai Zhang, Yi Luan, Hexiang Hu, Kenton Lee 외

Image retrieval, i.e., finding desired images given a reference image, inherently encompasses rich, multi-faceted search intents that are difficult to capture solely using image-based measures. Recent works leverage text…

Image RetrievalImplicit RelationsRetrievalSupervised Image Retrieval+2

M2-Encoder: Advancing Bilingual Image-Text Understanding by Large-scale Efficient Pretraining

2024-01-29 · Qingpei Guo, Furong Xu, Hanxiao Zhang, Wang Ren 외

Vision-language foundation models like CLIP have revolutionized the field of artificial intelligence. Nevertheless, VLM models supporting multi-language, e.g., in both Chinese and English, have lagged due to the relative…

GPUzero-shot-classificationZero-Shot Cross-Modal RetrievalZero-shot Image Retrieval+3

Linguistic-Aware Patch Slimming Framework for Fine-grained Cross-Modal Alignment

2024-01-01 · CVPR 2024 1 · Zheren Fu, Lei Zhang, Hou Xia, Zhendong Mao

Cross-modal alignment aims to build a bridge connecting vision and language. It is an important multi-modal task that efficiently learns the semantic similarities between images and texts. Traditional fine-grained al…

cross-modal alignmentCross-Modal RetrievalImage RetrievalImage-to-Text Retrieval+4

ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities

2023-05-18 · Peng Wang, Shijie Wang, Junyang Lin, Shuai Bai 외

In this work, we explore a scalable way for building a general representation model toward unlimited modalities. We release ONE-PEACE, a highly extensible model with 4B parameters that can seamlessly align and integrate …

1 Image, 2*2 StitchiAction ClassificationAudioCapsAudio Classification+18

CAVL: Learning Contrastive and Adaptive Representations of Vision and Language

2023-04-10 · Shentong Mo, Jingfei Xia, Ihor Markevych

Visual and linguistic pre-training aims to learn vision and language representations together, which can be transferred to visual-linguistic downstream tasks. However, there exists semantic confusion between language and…

Image RetrievalPhrase GroundingQuestion AnsweringRetrieval+6

Sigmoid Loss for Language Image Pre-Training

2023-03-27 · ICCV 2023 1 · Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas Beyer

We propose a simple pairwise Sigmoid loss for Language-Image Pre-training (SigLIP). Unlike standard contrastive learning with softmax normalization, the sigmoid loss operates solely on image-text pairs and does not requi…

Contrastive LearningDisentanglementImage-to-Text RetrievalZero-shot Text-to-Image Retrieval

BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

2023-01-30 · Conference 2023 2 · Junnan Li, Dongxu Li, Silvio Savarese, Steven Hoi

The cost of vision-and-language pre-training has become increasingly prohibitive due to end-to-end training of large-scale models. This paper proposes BLIP-2, a generic and efficient pre-training strategy that bootstraps…

Generative Visual Question AnsweringImage CaptioningImage RetrievalImage to text+13

Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

2022-11-02 · An Yang, Junshu Pan, Junyang Lin, Rui Men 외

The tremendous success of CLIP (Radford et al., 2021) has promoted the research and application of contrastive learning for vision-language pretraining. In this work, we construct a large-scale dataset of image-text pair…

Contrastive Learningimage-classificationImage ClassificationImage Retrieval+7

ERNIE-ViL 2.0: Multi-view Contrastive Learning for Image-Text Pre-training

2022-09-30 · Bin Shan, Weichong Yin, Yu Sun, Hao Tian 외

Recent Vision-Language Pre-trained (VLP) models based on dual encoder have attracted extensive attention from academia and industry due to their superior performance on various cross-modal tasks and high computational ef…

Computational EfficiencyContrastive LearningCross-Modal RetrievalImage Retrieval+5

Crossmodal-3600: A Massively Multilingual Multimodal Evaluation Dataset

2022-05-25 · Ashish V. Thapliyal, Jordi Pont-Tuset, Xi Chen, Radu Soricut

Research in massively multilingual image captioning has been severely hampered by a lack of high-quality evaluation datasets. In this paper we present the Crossmodal-3600 dataset (XM3600 in short), a geographically diver…

Image CaptioningImage RetrievalImage-text RetrievalImage-to-Text Retrieval+3

FLAVA: A Foundational Language And Vision Alignment Model

2021-12-08 · CVPR 2022 1 · Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon 외

State-of-the-art vision and vision-and-language models rely on large-scale visio-linguistic pretraining for obtaining good performance on a variety of downstream tasks. Generally, such models are often either cross-modal…

Image RetrievalImage-to-Text RetrievalVisual ReasoningZero-shot Image Retrieval+2

Learning Transferable Visual Models From Natural Language Supervision

2021-02-26 · Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 외

State-of-the-art computer vision systems are trained to predict a fixed set of predetermined object categories. This restricted form of supervision limits their generality and usability since additional labeled data is n…

Action RecognitionBenchmarkingFew-Shot Image Classificationgeo-localization+21

ZSCRGAN: A GAN-based Expectation Maximization Model for Zero-Shot Retrieval of Images from Textual Descriptions

2020-07-23 · Anurag Roy, Vinay Kumar Verma, Kripabandhu Ghosh, Saptarshi Ghosh

Most existing algorithms for cross-modal Information Retrieval are based on a supervised train-test setup, where a model learns to align the mode of the query (e.g., text) to the mode of the documents (e.g., images) from…

Cross-Modal Information RetrievalImage RetrievalInformation RetrievalRetrieval+3
1–15 / 15