paper-with-me

Papers

Document Haystack: A Long Context Multimodal Image/Document Understanding Vision LLM Benchmark

2025-07-18 · Goeric Huybrechts, Srikanth Ronanki, Sai Muralidhar Jayanthi, Jack Fitzgerald, Srinivasan Veeravanallur arxiv

The proliferation of multimodal Large Language Models has significantly advanced the ability to analyze and understand complex data inputs from different modalities. However, the processing of long documents remains under-explored, largely due to a lack of suitable benchmarks. To address this, we introduce Document Haystack, a comprehensive benchmark designed to evaluate the performance of Vision Language Models (VLMs) on long, visually complex documents. Document Haystack features documents ranging from 5 to 200 pages and strategically inserts pure text or multimodal text+image "needles" at various depths within the documents to challenge VLMs' retrieval capabilities. Comprising 400 document variants and a total of 8,250 questions, it is supported by an objective, automated evaluation framework. We detail the construction and characteristics of the Document Haystack dataset, present results from prominent VLMs and discuss potential research avenues in this area.

📄 PDF Abstract BibTeX arXiv:2507.15882

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Multimodal Needle in a Haystack: Benchmarking Long-Context Capability of Multimodal Large Language Models

2024-06-17 · Hengyi Wang, Haizhou Shi, Shiwei Tan, Weiyi Qin 외

Multimodal Large Language Models (MLLMs) have shown significant promise in various applications, leading to broad interest from researchers and practitioners alike. However, a comprehensive evaluation of their long-conte…

BenchmarkingHallucinationImage Retrieval+4

Document Haystacks: Vision-Language Reasoning Over Piles of 1000+ Documents

2024-11-23 · CVPR 2025 1 · Jun Chen, Dannong Xu, Junjie Fei, Chun-Mei Feng 외

Large multimodal models (LMMs) have achieved impressive progress in vision-language understanding, yet they face limitations in real-world applications requiring complex reasoning over a large number of images. Existing …

Question AnsweringRAGRetrievalRetrieval-augmented Generation+1

Summary of a Haystack: A Challenge to Long-Context LLMs and RAG Systems

2024-07-01 · Philippe Laban, Alexander R. Fabbri, Caiming Xiong, Chien-Sheng Wu

LLMs and RAG systems are now capable of handling millions of input tokens or more. However, evaluating the output quality of such systems on long-context tasks remains challenging, as tasks like Needle-in-a-Haystack lack…

RAG

Needle In A Multimodal Haystack

2024-06-11 · Weiyun Wang, Shuibo Zhang, Yiming Ren, Yuchen Duan 외

With the rapid advancement of multimodal large language models (MLLMs), their evaluation has become increasingly comprehensive. However, understanding long multimodal content, as a foundational ability for real-world app…

Retrieval

MultiHaystack: Benchmarking Multimodal Retrieval and Reasoning over 40K Images, Videos, and Documents

2026-03-05 · Dannong Xu, Zhongyu Yang, Jun Chen, Yingfang Yuan 외 arxiv

Multimodal large language models (MLLMs) achieve strong performance on benchmarks that evaluate text, image, or video understanding separately. However, these settings do not assess a critical real-world requirement, whi…