paper-with-me

Papers

Towards Natural Language-Based Document Image Retrieval: New Dataset and Benchmark

2025-01-01 · CVPR 2025 1 · Hao Guo, Xugong Qin, Jun Jie Ou Yang, Peng Zhang, Gangyan Zeng, Yubo Li, Hailun Lin

Document image retrieval (DIR) aims to retrieve document images from a gallery according to a given query. Existing DIR methods are primarily based on image queries that retrieve documents within the same coarse semantic category, e.g., newspapers or receipts. However, these methods struggle to effectively retrieve document images in real-world scenarios where textual queries with fine-grained semantics are usually provided. To bridge this gap, we introduce a new Natural Language-based Document Image Retrieval (NL-DIR) benchmark with corresponding evaluation metrics. In this work, natural language descriptions serve as semantically rich queries for the DIR task. The NL-DIR dataset contains 41K authentic document images, each paired with five high-quality, fine-grained semantic queries generated and evaluated through large language models in conjunction with manual verification. We perform zero-shot and fine-tuning evaluations of existing mainstream contrastive vision-language models and OCR-free visual document understanding (VDU) models. A two-stage retrieval method is further investigated for performance improvement while achieving both time and space efficiency. We hope the proposed NL-DIR benchmark can bring new opportunities and facilitate research for the VDU community. Datasets and codes will be publicly available at huggingface.co/datasets/nianbing/NL-DIR.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

document understandingImage RetrievalOptical Character Recognition (OCR)Retrieval

Similar Papers 제목 키워드 기반

Towards Natural Language-Based Document Image Retrieval: New Dataset and Benchmark

2025-12-23 · Hao Guo, Xugong Qin, Jun Jie Ou Yang, Peng Zhang 외 arxiv

Document image retrieval (DIR) aims to retrieve document images from a gallery according to a given query. Existing DIR methods are primarily based on image queries that retrieve documents within the same coarse semantic…

Image Retrieval

Unifying Multimodal Retrieval via Document Screenshot Embedding

2024-06-17 · Xueguang Ma, Sheng-Chieh Lin, Minghan Li, Wenhu Chen 외

In the real world, documents are organized in different formats and varied modalities. Traditional retrieval pipelines require tailored document parsing techniques and content extraction modules to prepare input for inde…

Language ModellingNatural QuestionsOptical Character Recognition (OCR)Retrieval+1

Design Challenges for a Multi-Perspective Search Engine

2021-12-15 · Findings (NAACL) 2022 7 · Sihao Chen, Siyi Liu, Xander Uyttendaele, Yi Zhang 외

Many users turn to document retrieval systems (e.g. search engines) to seek answers to controversial questions. Answering such user queries usually require identifying responses within web documents, and aggregating the …

Natural Language UnderstandingRetrieval

MIRACL-VISION: A Large, multilingual, visual document retrieval benchmark

2025-05-16 · Radek Osmulski, Gabriel de Souza P. Moreira, Ronay Ak, Mengyao Xu 외

Document retrieval is an important task for search and Retrieval-Augmented Generation (RAG) applications. Large Language Models (LLMs) have contributed to improving the accuracy of text-based document retrieval. However,…

RAGRetrievalRetrieval-augmented Generation

Audio Retrieval for Multimodal Design Documents: A New Dataset and Algorithms

2023-02-28 · Prachi Singh, Srikrishna Karanam, Sumit Shekhar

We consider and propose a new problem of retrieving audio files relevant to multimodal design document inputs comprising both textual elements and visual imagery, e.g., birthday/greeting cards. In addition to enhancing u…

Retrieval