paper-with-me

Image-text matching

1개 벤치마크 · 논문 214편 · 이 태스크의 논문 보기 →

Benchmarks

CommercialAdsDataset

결과 24개

Most implemented

Papers

Variational Adapter for Cross-modal Similarity Representation

2026-05-29 · WenZhang Wei, Zhipeng Gui, Dehua Peng, Tiandi Ye 외 arxiv

The core of vision-language models lies in measuring cross-modal similarity within a unified representation space. However, most image-text matching or multi-class image classification datasets lack fine-grained cross-mo…

Domain GeneralizationBinary ClassificationImage ClassificationImage-text matching

TempRet: Temporal Enhancement and Two-Stage Reranking for CVPR 2026 EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge

2026-05-23 · Zixu Li, Yupeng Hu, Zhiwei Chen, Zhiheng Fu 외 arxiv

Video-text retrieval has witnessed remarkable progress driven by large-scale vision-language pretraining, yet most existing approaches inherit an implicit assumption from image-text retrieval: that visual semantics can b…

Video-Text RetrievalImage-text matchingVideo Retrieval

MSD-Score: Multi-Scale Distributional Scoring for Reference-Free Image Caption Evaluation

2026-05-07 · Shichao Kan, Xuyang Zhang, Haojie Zhang, Zhe Zhu 외 arxiv

Evaluating image captions without references remains challenging because global embedding similarity often misses fine-grained mismatches such as hallucinated objects, missing attributes, or incorrect relations. We propo…

Image-text matching

Harnessing Weak Pair Uncertainty for Text-based Person Search

2026-04-10 · Jintao Sun, Zhedong Zheng, Gangyi Ding arxiv

In this paper, we study the text-based person search, which is to retrieve the person of interest via natural language description. Prevailing methods usually focus on the strict one-to-one correspondence pair matching b…

Contrastive LearningImage-text matchingPerson Search

DBMF: A Dual-Branch Multimodal Framework for Out-of-Distribution Detection

2026-04-09 · Jiangbei Yue, Darren Treanor, Venkataraman Subramanian, Sharib Ali arxiv

The complex and dynamic real-world clinical environment demands reliable deep learning (DL) systems. Out-of-distribution (OOD) detection plays a critical role in enhancing the reliability and generalizability of DL model…

Out-of-Distribution DetectionImage-text matching

Hidden in the Multiplicative Interaction: Uncovering Fragility in Multimodal Contrastive Learning

2026-04-07 · Tillmann Rheude, Stefan Hegselmann, Roland Eils, Benjamin Wild arxiv

Contrastive learning has become a standard approach for unsupervised learning from paired data, as demonstrated by CLIP for image-text matching. However, many domains involve more than two modalities and require objectiv…

Cross-Modal RetrievalContrastive LearningImage-text matching

전체 214편 보기 →