paper-with-me

Text Retrieval

16개 벤치마크 · 논문 788편 · 이 태스크의 논문 보기 →

Benchmarks

MTEB

결과 30개

20 Newsgroups

결과 3개

Image-Chat

결과 3개

Reuters-21578

결과 3개

CLIMATE-FEVER

결과 1개

DBpedia

결과 1개

FEVER

결과 1개

HotpotQA

결과 1개

MS MARCO

결과 1개

NFCorpus

결과 1개

Natural Questions

결과 1개

Quora Question Pairs

결과 1개

RSICD

결과 1개

SciDocs

결과 1개

SciFact

결과 1개

TREC-COVID

결과 1개

Most implemented

FlexiViT: One Model for All Patch Sizes

2022-12-15 · 구현 6개

Papers

Retrieval Heads Meet Vision: Uncovering How VLMs Locate and Extract Visual Information

2026-08-27 · Chanho Park, Daehyeon Choi, Jihyun Lee, Minhyuk Sung arxiv

Vision-language models (VLMs) can locate an image region referred to by a text prompt and route the corresponding visual evidence to the output, yet the internal mechanism behind this behavior is not understood. Inspired…

Text Retrieval

MLLMCLIP: Feature-Level Distillation of MLLM for Robust Vision-Language Representations

2026-08-26 · Jongsuk Kim, Qiyu Wu, Zhuoyuan Mao, Hiromi Wakaki 외 arxiv

Pretrained vision-language models such as CLIP excel at zero-shot recognition but often fail at compositionality, particularly attribute-object and relational structures. Recent studies mitigate this issue by augmenting …

Text Retrieval

LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training

2026-08-25 · Andreas Hochlehnert, Marianna Nezhurina, Mehdi Cherti, Andrej Radonjic 외 hf

We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from CommonCrawl. From these, we download 80M videos with a total duration of 10 …

Text Retrieval

ConceptFormer: Learning Adaptive Latent Concepts for Query-Document Alignment in Visual Document Retrieval

2026-08-16 · Peng Chunyi, Xu Zhipeng, Yan Yukun, Liu Zhenghao 외 hf

Visual document retrieval is a critical component of multimodal retrieval-augmented generation, aiming to identify query-relevant pages from document collections where evidence is distributed across text, layout, charts,…

Representation LearningText Retrieval

DARAD: Dual Adapters and Ranking-Aware Distillation for Continual Remote Sensing Image-Text Retrieval

2026-08-06 · Xi Chen, Xu Chen, Xiangyang Jia, Wei Wang 외 arxiv

With the rapid growth of Earth observation technologies, remote sensing archives are rapidly expanding, making remote sensing image-text retrieval (RS-ITR) increasingly important. However, continual RS-ITR remains challe…

Continual LearningText Retrieval

Through the LENS: Local Geometric Decomposition of Vision-Language Model Representations

2026-08-01 · Shalom Kachko, Raz Lapid, Margarita Vald, Almog Dubin 외 arxiv

Vision-language models (VLMs) process image patches and text tokens in a shared residual stream, but the local geometry through which the two modalities interact remains poorly understood. Most interpretability methods i…

Text Retrieval

전체 788편 보기 →