paper-with-me

Papers Image-text Retrieval

“Image-text Retrieval” 태그가 달린 논문 248편 · 필터 해제

Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal Retrieval

2025-06-26 · Hani AlOmari, Anushka Sivakumar, Andrew Zhang, Chris Thomas

Cross-modal image-text retrieval is challenging because of the diverse possible associations between content from different modalities. Traditional methods learn a single-vector embedding to represent semantics of each s…

Cross-Modal RetrievalImage-text RetrievalRetrievalText Retrieval

Adding simple structure at inference improves Vision-Language Compositionality

2025-06-11 · Imanol Miranda, Ander Salaberria, Eneko Agirre, Gorka Azkune

Dual encoder Vision-Language Models (VLM) such as CLIP are widely used for image-text retrieval tasks. However, those models struggle with compositionality, showing a bag-of-words-like behavior that limits their retrieva…

AttributeImage-text RetrievalRetrievalText Retrieval+1

FlagEvalMM: A Flexible Framework for Comprehensive Multimodal Model Evaluation

2025-06-10 · Zheqi He, Yesheng Liu, Jing-shu Zheng, Xuejing Li 외

We present FlagEvalMM, an open-source evaluation framework designed to comprehensively assess multimodal models across a diverse range of vision-language understanding and generation tasks, such as visual question answer…

Image-text RetrievalQuestion AnsweringText RetrievalVideo Generation+1

Attacking Attention of Foundation Models Disrupts Downstream Tasks

2025-06-03 · Hondamunige Prasanna Silva, Federico Becattini, Lorenzo Seidenari

Foundation models represent the most prominent and recent paradigm shift in artificial intelligence. Foundation models are large models, trained on broad data that deliver high accuracy in many downstream tasks, often wi…

Depth EstimationImage-text RetrievalText Retrieval

Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation

2025-05-25 · Daniel Csizmadia, Andrei Codreanu, Victor Sim, Vighnesh Prabhu 외

We present Distill CLIP (DCLIP), a fine-tuned variant of the CLIP model that enhances multimodal image-text retrieval while preserving the original model's strong zero-shot classification capabilities. CLIP models are ty…

Contrastive LearningImage-text RetrievalRetrievalText Retrieval+2

EvdCLIP: Improving Vision-Language Retrieval with Entity Visual Descriptions from Large Language Models

2025-05-24 · Guanghao Meng, Sunan He, Jinpeng Wang, Tao Dai 외

Vision-language retrieval (VLR) has attracted significant attention in both academia and industry, which involves using text (or images) as queries to retrieve corresponding images (or text). However, existing methods of…

Image-text RetrievalLanguage ModelingLanguage ModellingLarge Language Model+2

Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval

2025-05-22 · Hailong Ning, Siying Wang, Tao Lei, Xiaopeng Cao 외

Remote Sensing Image-Text Retrieval (RSITR) plays a critical role in geographic information interpretation, disaster monitoring, and urban planning by establishing semantic associations between image and textual descript…

cross-modal alignmentImage-text Retrievalparameter-efficient fine-tuningRetrieval+1

Breaking Language Barriers or Reinforcing Bias? A Study of Gender and Racial Disparities in Multilingual Contrastive Vision Language Models

2025-05-20 · Zahraa Al Sahili, Ioannis Patras, Matthew Purver

Multilingual vision-language models promise universal image-text retrieval, yet their social biases remain under-explored. We present the first systematic audit of three public multilingual CLIP checkpoints -- M-CLIP, NL…

Image-text RetrievalText Retrieval

A Vision-Language Foundation Model for Leaf Disease Identification

2025-05-11 · Khang Nguyen Quoc, Lan Le Thi Thu, Luyl-Da Quach

Leaf disease identification plays a pivotal role in smart agriculture. However, many existing studies still struggle to integrate image and textual modalities to compensate for each other's limitations. Furthermore, many…

Contrastive Learningimage-classificationImage ClassificationImage-text Retrieval+1

FG-CLIP: Fine-Grained Visual and Textual Alignment

2025-05-08 · Chunyu Xie, Bin Wang, Fanjing Kong, Jincheng Li 외

Contrastive Language-Image Pre-training (CLIP) excels in multimodal tasks such as image-text retrieval and zero-shot classification but struggles with fine-grained understanding due to its focus on coarse-grained short c…

Image-text Retrievalobject-detectionObject DetectionOpen-vocabulary object detection+5

AGATE: Stealthy Black-box Watermarking for Multimodal Model Copyright Protection

2025-04-28 · Jianbo Gao, Keke Gai, Jing Yu, Liehuang Zhu 외

Recent advancement in large-scale Artificial Intelligence (AI) models offering multimodal services have become foundational in AI systems, making them prime targets for model theft. Existing methods select Out-of-Distrib…

Adversarial AttackAnomaly Detectionimage-classificationImage Classification+3

Breaking the Modality Barrier: Universal Embedding Learning with Multimodal LLMs

2025-04-24 · Tiancheng Gu, Kaicheng Yang, Ziyong Feng, Xingjun Wang 외

The Contrastive Language-Image Pre-training (CLIP) framework has become a widely used approach for multimodal representation learning, particularly in image-text retrieval and clustering. However, its efficacy is constra…

Image-text RetrievalInstruction FollowingKnowledge DistillationRepresentation Learning+2

FocalLens: Instruction Tuning Enables Zero-Shot Conditional Image Representations

2025-04-11 · Cheng-Yu Hsieh, Pavan Kumar Anasosalu Vasu, Fartash Faghri, Raviteja Vemulapalli 외

Visual understanding is inherently contextual -- what we focus on in an image depends on the task at hand. For instance, given an image of a person holding a bouquet of flowers, we may focus on either the person such as …

image-classificationImage ClassificationImage RetrievalImage-text Retrieval+2

Med3DVLM: An Efficient Vision-Language Model for 3D Medical Image Analysis

2025-03-25 · Yu Xin, Gorkem Can Ates, Kuang Gong, Wei Shao

Vision-language models (VLMs) have shown promise in 2D medical image analysis, but extending them to 3D remains challenging due to the high computational demands of volumetric data and the difficulty of aligning 3D spati…

Contrastive LearningImage-text RetrievalLanguage ModelingLanguage Modelling+5

SeLIP: Similarity Enhanced Contrastive Language Image Pretraining for Multi-modal Head MRI

2025-03-25 · Zhiyang Liu, Dong Yang, Minghao Zhang, Hanyu Sun 외

Despite that deep learning (DL) methods have presented tremendous potential in many medical image analysis tasks, the practical applications of medical DL models are limited due to the lack of enough data samples with ma…

Contrastive LearningImage SegmentationImage-text RetrievalMedical Image Analysis+4

Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

2025-03-25 · Ilias Stogiannidis, Steven McDonagh, Sotirios A. Tsaftaris

Vision-Language Models (VLMs) have recently emerged as powerful tools, excelling in tasks that integrate visual and textual comprehension, such as image captioning, visual question answering, and image-text retrieval. Ho…

BenchmarkingImage CaptioningImage-text Retrievalobject-detection+5

Anatomy-Aware Conditional Image-Text Retrieval

2025-03-10 · Meng Zheng, Jiajin Zhang, Benjamin Planche, Zhongpai Gao 외

Image-Text Retrieval (ITR) finds broad applications in healthcare, aiding clinicians and radiologists by automatically retrieving relevant patient cases in the database given the query image and/or report, for more effic…

AnatomyContrastive LearningImage-text RetrievalRetrieval+1

Variance-Aware Loss Scheduling for Multimodal Alignment in Low-Data Settings

2025-03-05 · Sneh Pillai

Training vision-language models for image-text alignment typically requires large datasets to achieve robust performance. In low-data scenarios, standard contrastive learning can struggle to align modalities effectively …

Contrastive LearningImage-text RetrievalSchedulingText Retrieval

LLaVE: Large Language and Vision Embedding Models with Hardness-Weighted Contrastive Learning

2025-03-04 · Zhibin Lan, LiQiang Niu, Fandong Meng, Jie zhou 외

Universal multimodal embedding models play a critical role in tasks such as interleaved image-text retrieval, multimodal RAG, and multimodal clustering. However, our empirical results indicate that existing LMM-based emb…

Contrastive LearningImage-text RetrievalRAGRepresentation Learning+3

MedUnifier: Unifying Vision-and-Language Pre-training on Medical Data with Vision Generation Task using Discrete Visual Representations

2025-03-02 · CVPR 2025 1 · Ziyang Zhang, Yang Yu, Yucheng Chen, Xulei Yang 외

Despite significant progress in Vision-Language Pre-training (VLP), current approaches predominantly emphasize feature extraction and cross-modal comprehension, with limited attention to generating or transforming visual…

image-classificationImage ClassificationImage GenerationImage-text matching+7
1–20 / 248 다음 →