Image-text matching
1개 벤치마크 · 논문 214편 · 이 태스크의 논문 보기 →
Benchmarks
CommercialAdsDataset
Most implemented
AttnGAN: Fine-Grained Text to Image Generation with Attentional Generative Adversarial Networks
BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
VinVL: Revisiting Visual Representations in Vision-Language Models
UNITER: UNiversal Image-TExt Representation Learning
Align before Fuse: Vision and Language Representation Learning with Momentum Distillation
Stacked Cross Attention for Image-Text Matching
Papers
Variational Adapter for Cross-modal Similarity Representation
The core of vision-language models lies in measuring cross-modal similarity within a unified representation space. However, most image-text matching or multi-class image classification datasets lack fine-grained cross-mo…
Domain GeneralizationBinary ClassificationImage ClassificationImage-text matchingTempRet: Temporal Enhancement and Two-Stage Reranking for CVPR 2026 EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge
Video-text retrieval has witnessed remarkable progress driven by large-scale vision-language pretraining, yet most existing approaches inherit an implicit assumption from image-text retrieval: that visual semantics can b…
Video-Text RetrievalImage-text matchingVideo RetrievalMSD-Score: Multi-Scale Distributional Scoring for Reference-Free Image Caption Evaluation
Evaluating image captions without references remains challenging because global embedding similarity often misses fine-grained mismatches such as hallucinated objects, missing attributes, or incorrect relations. We propo…
Image-text matchingHarnessing Weak Pair Uncertainty for Text-based Person Search
In this paper, we study the text-based person search, which is to retrieve the person of interest via natural language description. Prevailing methods usually focus on the strict one-to-one correspondence pair matching b…
Contrastive LearningImage-text matchingPerson SearchDBMF: A Dual-Branch Multimodal Framework for Out-of-Distribution Detection
The complex and dynamic real-world clinical environment demands reliable deep learning (DL) systems. Out-of-distribution (OOD) detection plays a critical role in enhancing the reliability and generalizability of DL model…
Out-of-Distribution DetectionImage-text matchingHidden in the Multiplicative Interaction: Uncovering Fragility in Multimodal Contrastive Learning
Contrastive learning has become a standard approach for unsupervised learning from paired data, as demonstrated by CLIP for image-text matching. However, many domains involve more than two modalities and require objectiv…
Cross-Modal RetrievalContrastive LearningImage-text matching