Zero-Shot Cross-Modal Retrieval
3개 벤치마크 · 논문 28편 · 이 태스크의 논문 보기 →
Benchmarks
Most implemented
Learning Transferable Visual Models From Natural Language Supervision
UNITER: UNiversal Image-TExt Representation Learning
CoCa: Contrastive Captioners are Image-Text Foundation Models
Align before Fuse: Vision and Language Representation Learning with Momentum Distillation
Reproducible scaling laws for contrastive language-image learning
Papers
MatBind: A Shared Embedding Space for Multimodal Materials Characterization
Fully characterizing a crystalline material requires integrating heterogeneous data sources -- atomic structures, diffraction patterns, electronic density of states, and natural language -- each of which captures a diffe…
Zero-Shot Cross-Modal RetrievalContrastive LearningSTAR: Semantic-Traffic Alignment and Retrieval for Zero-Shot HTTPS Website Fingerprinting
Modern HTTPS mechanisms such as Encrypted Client Hello (ECH) and encrypted DNS improve privacy but remain vulnerable to website fingerprinting (WF) attacks, where adversaries infer visited sites from encrypted traffic pa…
Zero-Shot Cross-Modal RetrievalFineLIP: Extending CLIP's Reach via Fine-Grained Alignment with Longer Text Inputs
As a pioneering vision-language model, CLIP (Contrastive Language-Image Pre-training) has achieved significant success across various domains and a wide range of downstream vision-language tasks. However, the text encode…
cross-modal alignmentCross-Modal RetrievalImage GenerationText to Image Generation+2A Recipe for Improving Remote Sensing VLM Zero Shot Generalization
Foundation models have had a significant impact across various AI applications, enabling use cases that were previously impossible. Contrastive Visual Language Models (VLMs), in particular, have outperformed other techni…
Cross-Modal RetrievalZero-Shot Cross-Modal RetrievalZero-shot GeneralizationIMPACT: A Large-scale Integrated Multimodal Patent Analysis and Creation Dataset for Design Patents
In this paper, we introduce IMPACT (Integrated Multimodal Patent Analysis and Creation Dataset for Design Patents), a large-scale multimodal patent dataset with detailed captions for design patent figures. Our dataset in…
Cross-Modal RetrievalImage ClassificationImage RetrievalPatent classification+5COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training
Vision-Language Models (VLMs) trained with contrastive loss have achieved significant advancements in various vision and language tasks. However, the global nature of the contrastive loss makes VLMs focus predominantly o…
Self-Supervised LearningSemantic SegmentationUnsupervised Semantic Segmentation with Language-image Pre-trainingZero-Shot Cross-Modal Retrieval+1