Text Retrieval
16개 벤치마크 · 논문 788편 · 이 태스크의 논문 보기 →
Benchmarks
MTEB
20 Newsgroups
Image-Chat
Reuters-21578
CLIMATE-FEVER
DBpedia
FEVER
HotpotQA
MS MARCO
NFCorpus
Natural Questions
Quora Question Pairs
RSICD
SciDocs
SciFact
TREC-COVID
Most implemented
BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
UNITER: UNiversal Image-TExt Representation Learning
LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment
FlexiViT: One Model for All Patch Sizes
Papers
Retrieval Heads Meet Vision: Uncovering How VLMs Locate and Extract Visual Information
Vision-language models (VLMs) can locate an image region referred to by a text prompt and route the corresponding visual evidence to the output, yet the internal mechanism behind this behavior is not understood. Inspired…
Text RetrievalMLLMCLIP: Feature-Level Distillation of MLLM for Robust Vision-Language Representations
Pretrained vision-language models such as CLIP excel at zero-shot recognition but often fail at compositionality, particularly attribute-object and relational structures. Recent studies mitigate this issue by augmenting …
Text RetrievalLAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training
We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from CommonCrawl. From these, we download 80M videos with a total duration of 10 …
Text RetrievalConceptFormer: Learning Adaptive Latent Concepts for Query-Document Alignment in Visual Document Retrieval
Visual document retrieval is a critical component of multimodal retrieval-augmented generation, aiming to identify query-relevant pages from document collections where evidence is distributed across text, layout, charts,…
Representation LearningText RetrievalDARAD: Dual Adapters and Ranking-Aware Distillation for Continual Remote Sensing Image-Text Retrieval
With the rapid growth of Earth observation technologies, remote sensing archives are rapidly expanding, making remote sensing image-text retrieval (RS-ITR) increasingly important. However, continual RS-ITR remains challe…
Continual LearningText RetrievalThrough the LENS: Local Geometric Decomposition of Vision-Language Model Representations
Vision-language models (VLMs) process image patches and text tokens in a shared residual stream, but the local geometry through which the two modalities interact remains poorly understood. Most interpretability methods i…
Text Retrieval