paper-with-me

홈 › Papers

XL$^2$Bench: A Benchmark for Extremely Long Context Understanding with Long-range Dependencies

2024-04-08 · Xuanfan Ni, Hengyi Cai, Xiaochi Wei, Shuaiqiang Wang, Dawei Yin, Piji Li

Large Language Models (LLMs) have demonstrated remarkable performance across diverse tasks but are constrained by their small context window sizes. Various efforts have been proposed to expand the context window to accommodate even up to 200K input tokens. Meanwhile, building high-quality benchmarks with much longer text lengths and more demanding tasks to provide comprehensive evaluations is of immense practical interest to facilitate long context understanding research of LLMs. However, prior benchmarks create datasets that ostensibly cater to long-text comprehension by expanding the input of traditional tasks, which falls short to exhibit the unique characteristics of long-text understanding, including long dependency tasks and longer text length compatible with modern LLMs' context window size. In this paper, we introduce a benchmark for extremely long context understanding with long-range dependencies, XL$^2$Bench, which includes three scenarios: Fiction Reading, Paper Reading, and Law Reading, and four tasks of increasing complexity: Memory Retrieval, Detailed Understanding, Overall Understanding, and Open-ended Generation, covering 27 subtasks in English and Chinese. It has an average length of 100K+ words (English) and 200K+ characters (Chinese). Evaluating six leading LLMs on XL$^2$Bench, we find that their performance significantly lags behind human levels. Moreover, the observed decline in performance across both the original and enhanced datasets underscores the efficacy of our approach to mitigating data contamination.

📄 PDF Abstract BibTeX arXiv:2404.05446

Code (0)

등록된 구현이 없습니다.

Tasks

Long-Context UnderstandingReading Comprehension

Similar Papers 제목 키워드 기반

X-LeBench: A Benchmark for Extremely Long Egocentric Video Understanding

2025-01-12 · Wenqi Zhou, Kai Cao, Hao Zheng, Xinyi Zheng 외

Long-form egocentric video understanding provides rich contextual information and unique insights into long-term human behaviors, holding significant potential for applications in embodied intelligence, long-term activit…

Video Understanding

NovelQA: Benchmarking Question Answering on Documents Exceeding 200K Tokens

2024-03-18 · Cunxiang Wang, Ruoxi Ning, Boqi Pan, Tonghui Wu 외

The rapid advancement of Large Language Models (LLMs) has introduced a new frontier in natural language processing, particularly in understanding and processing long-context information. However, the evaluation of these …

BenchmarkingQuestion Answering

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

2025-06-30 · Haoji Zhang, Yiqin Wang, Yansong Tang, Yong liu 외

Benefiting from the advances in large language models and cross-modal alignment, existing multimodal large language models have achieved prominent performance in image and short video understanding. However, the understa…

cross-modal alignmentEgoSchemaMMEMVBench+2

Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks

2024-04-09 · Chonghua Wang, Haodong Duan, Songyang Zhang, Dahua Lin 외

Recently, the large language model (LLM) community has shown increasing interest in enhancing LLMs' capability to handle extremely long documents. As various long-text techniques and model architectures emerge, the preci…

Answer SelectionLong-Context UnderstandingQuestion Answering

MedHorizon: Towards Long-context Medical Video Understanding in the Wild

2026-05-07 · Bodong Du, Bowen Liu, Yang Yu, Xinpeng Ding 외 arxiv

Medical multimodal large language models (MLLMs) have advanced image understanding and short-video analysis, but real clinical review often requires full-procedure video understanding. Unlike general long videos, medical…