Image-to-Text Retrieval
8개 벤치마크 · 논문 66편 · 이 태스크의 논문 보기 →
Benchmarks
Flickr30k
WHOOPS!
AIC-ICC
COCO
FETA Car-Manuals
RSICD
RUC-CAS-WenLan
Most implemented
Learning Transferable Visual Models From Natural Language Supervision
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Sigmoid Loss for Language Image Pre-Training
Align before Fuse: Vision and Language Representation Learning with Momentum Distillation
FLAVA: A Foundational Language And Vision Alignment Model
Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks
Papers
OrganLens: Organ-Specific Representation Learning for CT Foundation Models
A CT examination captures multiple organs, but many biomedical questions concern abnormalities, prognosis, or longitudinal change in a specific organ. These questions require a separate representation for each organ with…
Representation LearningImage-to-Text RetrievalOne Single Hub Text Breaks CLIP: Identifying Vulnerabilities in Cross-Modal Encoders via Hubness
The hubness problem, in which hub embeddings are close to many unrelated examples, occurs often in high-dimensional embedding spaces and may pose a practical threat for purposes such as information retrieval and automati…
Image-to-Text RetrievalInformation RetrievalImage CaptioningNegative Entity Suppression for Zero-Shot Captioning with Synthetic Images
Text-only training provides an attractive approach to address data scarcity challenges in zero-shot image captioning (ZIC), avoiding the expense of collecting paired image-text annotations. However, although these approa…
Image-to-Text RetrievalDomain GeneralizationImage CaptioningEvaluating Perspectival Biases in Cross-Modal Retrieval
Multimodal retrieval systems are expected to operate in a semantic space, agnostic to the language or cultural origin of the query. In practice, however, retrieval outcomes systematically reflect perspectival biases: dev…
Representation LearningImage-to-Text RetrievalCross-Modal RetrievalImage RetrievalDualCap: Enhancing Lightweight Image Captioning via Dual Retrieval with Similar Scenes Visual Prompts
Recent lightweight retrieval-augmented image caption models often utilize retrieved data solely as text prompts, thereby creating a semantic gap by leaving the original visual features unenhanced, particularly for object…
Image-to-Text RetrievalImage CaptioningImage RetrievalMeta CLIP 2: A Worldwide Scaling Recipe
Contrastive Language-Image Pretraining (CLIP) is a popular foundation model, supporting from zero-shot classification, retrieval to encoders for multimodal large language models (MLLMs). Although CLIP is successfully tra…
Image-to-Text Retrieval