Papers Image-text Retrieval
“Image-text Retrieval” 태그가 달린 논문 248편 · 필터 해제
Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal Retrieval
Cross-modal image-text retrieval is challenging because of the diverse possible associations between content from different modalities. Traditional methods learn a single-vector embedding to represent semantics of each s…
Cross-Modal RetrievalImage-text RetrievalRetrievalText RetrievalAdding simple structure at inference improves Vision-Language Compositionality
Dual encoder Vision-Language Models (VLM) such as CLIP are widely used for image-text retrieval tasks. However, those models struggle with compositionality, showing a bag-of-words-like behavior that limits their retrieva…
AttributeImage-text RetrievalRetrievalText Retrieval+1FlagEvalMM: A Flexible Framework for Comprehensive Multimodal Model Evaluation
We present FlagEvalMM, an open-source evaluation framework designed to comprehensively assess multimodal models across a diverse range of vision-language understanding and generation tasks, such as visual question answer…
Image-text RetrievalQuestion AnsweringText RetrievalVideo Generation+1Attacking Attention of Foundation Models Disrupts Downstream Tasks
Foundation models represent the most prominent and recent paradigm shift in artificial intelligence. Foundation models are large models, trained on broad data that deliver high accuracy in many downstream tasks, often wi…
Depth EstimationImage-text RetrievalText RetrievalDistill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation
We present Distill CLIP (DCLIP), a fine-tuned variant of the CLIP model that enhances multimodal image-text retrieval while preserving the original model's strong zero-shot classification capabilities. CLIP models are ty…
Contrastive LearningImage-text RetrievalRetrievalText Retrieval+2EvdCLIP: Improving Vision-Language Retrieval with Entity Visual Descriptions from Large Language Models
Vision-language retrieval (VLR) has attracted significant attention in both academia and industry, which involves using text (or images) as queries to retrieve corresponding images (or text). However, existing methods of…
Image-text RetrievalLanguage ModelingLanguage ModellingLarge Language Model+2Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval
Remote Sensing Image-Text Retrieval (RSITR) plays a critical role in geographic information interpretation, disaster monitoring, and urban planning by establishing semantic associations between image and textual descript…
cross-modal alignmentImage-text Retrievalparameter-efficient fine-tuningRetrieval+1Breaking Language Barriers or Reinforcing Bias? A Study of Gender and Racial Disparities in Multilingual Contrastive Vision Language Models
Multilingual vision-language models promise universal image-text retrieval, yet their social biases remain under-explored. We present the first systematic audit of three public multilingual CLIP checkpoints -- M-CLIP, NL…
Image-text RetrievalText RetrievalA Vision-Language Foundation Model for Leaf Disease Identification
Leaf disease identification plays a pivotal role in smart agriculture. However, many existing studies still struggle to integrate image and textual modalities to compensate for each other's limitations. Furthermore, many…
Contrastive Learningimage-classificationImage ClassificationImage-text Retrieval+1FG-CLIP: Fine-Grained Visual and Textual Alignment
Contrastive Language-Image Pre-training (CLIP) excels in multimodal tasks such as image-text retrieval and zero-shot classification but struggles with fine-grained understanding due to its focus on coarse-grained short c…
Image-text Retrievalobject-detectionObject DetectionOpen-vocabulary object detection+5AGATE: Stealthy Black-box Watermarking for Multimodal Model Copyright Protection
Recent advancement in large-scale Artificial Intelligence (AI) models offering multimodal services have become foundational in AI systems, making them prime targets for model theft. Existing methods select Out-of-Distrib…
Adversarial AttackAnomaly Detectionimage-classificationImage Classification+3Breaking the Modality Barrier: Universal Embedding Learning with Multimodal LLMs
The Contrastive Language-Image Pre-training (CLIP) framework has become a widely used approach for multimodal representation learning, particularly in image-text retrieval and clustering. However, its efficacy is constra…
Image-text RetrievalInstruction FollowingKnowledge DistillationRepresentation Learning+2FocalLens: Instruction Tuning Enables Zero-Shot Conditional Image Representations
Visual understanding is inherently contextual -- what we focus on in an image depends on the task at hand. For instance, given an image of a person holding a bouquet of flowers, we may focus on either the person such as …
image-classificationImage ClassificationImage RetrievalImage-text Retrieval+2Med3DVLM: An Efficient Vision-Language Model for 3D Medical Image Analysis
Vision-language models (VLMs) have shown promise in 2D medical image analysis, but extending them to 3D remains challenging due to the high computational demands of volumetric data and the difficulty of aligning 3D spati…
Contrastive LearningImage-text RetrievalLanguage ModelingLanguage Modelling+5SeLIP: Similarity Enhanced Contrastive Language Image Pretraining for Multi-modal Head MRI
Despite that deep learning (DL) methods have presented tremendous potential in many medical image analysis tasks, the practical applications of medical DL models are limited due to the lack of enough data samples with ma…
Contrastive LearningImage SegmentationImage-text RetrievalMedical Image Analysis+4Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models
Vision-Language Models (VLMs) have recently emerged as powerful tools, excelling in tasks that integrate visual and textual comprehension, such as image captioning, visual question answering, and image-text retrieval. Ho…
BenchmarkingImage CaptioningImage-text Retrievalobject-detection+5Anatomy-Aware Conditional Image-Text Retrieval
Image-Text Retrieval (ITR) finds broad applications in healthcare, aiding clinicians and radiologists by automatically retrieving relevant patient cases in the database given the query image and/or report, for more effic…
AnatomyContrastive LearningImage-text RetrievalRetrieval+1Variance-Aware Loss Scheduling for Multimodal Alignment in Low-Data Settings
Training vision-language models for image-text alignment typically requires large datasets to achieve robust performance. In low-data scenarios, standard contrastive learning can struggle to align modalities effectively …
Contrastive LearningImage-text RetrievalSchedulingText RetrievalLLaVE: Large Language and Vision Embedding Models with Hardness-Weighted Contrastive Learning
Universal multimodal embedding models play a critical role in tasks such as interleaved image-text retrieval, multimodal RAG, and multimodal clustering. However, our empirical results indicate that existing LMM-based emb…
Contrastive LearningImage-text RetrievalRAGRepresentation Learning+3MedUnifier: Unifying Vision-and-Language Pre-training on Medical Data with Vision Generation Task using Discrete Visual Representations
Despite significant progress in Vision-Language Pre-training (VLP), current approaches predominantly emphasize feature extraction and cross-modal comprehension, with limited attention to generating or transforming visual…
image-classificationImage ClassificationImage GenerationImage-text matching+7