paper-with-me

홈 › Papers

PRELUDE: A Benchmark Designed to Require Global Comprehension and Reasoning over Long Contexts

2025-08-13 · Mo Yu, Tsz Ting Chung, Chulun Zhou, Tong Li, Rui Lu, Jiangnan Li, Liyan Xu, Haoshu Lu, Ning Zhang, Jing Li, Jie Zhou arxiv

We introduce PRELUDE, a benchmark for evaluating long-context understanding through the task of determining whether a character's prequel story is consistent with the canonical narrative of the original book. Our task poses a stronger demand for global comprehension and deep reasoning than existing benchmarks -- as the prequels are not part of the original story, assessing their plausibility typically requires searching and integrating information that is only indirectly related. Empirically, 88% of instances require evidence from multiple parts of the narrative. Experimental results highlight the challenge of our task: in-context learning, RAG and in-domain training with state-of-the-art LLMs, and commercial DeepResearch services, lag behind humans by >15%. A further human study reveals that models often produce correct answers with flawed reasoning, leading to an over 30% gap in reasoning accuracy compared to humans. These findings underscore the substantial room for improvement in long-context understanding and reasoning.

📄 PDF Abstract BibTeX arXiv:2508.09848

Code (0)

등록된 구현이 없습니다.

Tasks

Long-Context Understanding

Similar Papers 제목 키워드 기반

CASCADE: LLM-Powered JavaScript Deobfuscator at Google

2025-07-23 · Shan Jiang, Pranoy Kovuri, David Tao, Zhixun Tan arxiv

Software obfuscation, particularly prevalent in JavaScript, hinders code comprehension and analysis, posing significant challenges to software testing, static analysis, and malware detection. This paper introduces CASCAD…

Malware Detection

Improving Chemical Understanding of LLMs via SMILES Parsing

2025-05-22 · Yunhui Jang, Jaehyung Kim, Sungsoo Ahn

Large language models (LLMs) are increasingly recognized as powerful tools for scientific discovery, particularly in molecular science. A fundamental requirement for these models is the ability to accurately understand m…

Graph Matchingscientific discovery

SatireDecoder: Visual Cascaded Decoupling for Enhancing Satirical Image Comprehension

2025-11-29 · Yue Jiang, Haiwei Xue, Minghao Han, Mingcheng Li 외 arxiv

Satire, a form of artistic expression combining humor with implicit critique, holds significant social value by illuminating societal issues. Despite its cultural and societal significance, satire comprehension, particul…

CART: Context-Anchored Recurrent Transformer -- A Parameter-Efficient Architecture with Learned Stability

2026-05-31 · Chad A. Capps arxiv

We present CART (Context-Anchored Recurrent Transformer), a parameter-efficient language model that reuses a single shared core block R times across depth. Unlike prior looped transformers that recompute key-value tensor…

LVBench: An Extreme Long Video Understanding Benchmark

2024-06-12 · Weihan Wang, Zehai He, Wenyi Hong, Yean Cheng 외

Recent progress in multimodal large language models has markedly enhanced the understanding of short videos (typically under one minute), and several evaluation datasets have emerged accordingly. However, these advanceme…

Decision MakingVideo Understanding