paper-with-me

홈 › Papers

NarraBench: A Comprehensive Framework for Narrative Benchmarking

2025-10-10 · Sil Hamilton, Matthew Wilkens, Andrew Piper arxiv

We present NarraBench, a theory-informed taxonomy of narrative-understanding tasks, as well as an associated survey of 78 existing benchmarks in the area. We find significant need for new evaluations covering aspects of narrative understanding that are either overlooked in current work or are poorly aligned with existing metrics. Specifically, we estimate that only 27% of narrative tasks are well captured by existing benchmarks, and we note that some areas -- including narrative events, style, perspective, and revelation -- are nearly absent from current evaluations. We also note the need for increased development of benchmarks capable of assessing constitutively subjective and perspectival aspects of narrative, that is, aspects for which there is generally no single correct answer. Our taxonomy, survey, and methodology are of value to NLP researchers seeking to test LLM narrative understanding.

📄 PDF Abstract BibTeX arXiv:2510.09869

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Benchmarking LLMs on the Semantic Overlap Summarization Task

2024-02-26 · John Salvador, Naman Bansal, Mousumi Akter, Souvika Sarkar 외

Semantic Overlap Summarization (SOS) is a constrained multi-document summarization task, where the constraint is to capture the common/overlapping information between two alternative narratives. While recent advancements…

BenchmarkingDocument SummarizationMulti-Document Summarization

SeqBench: Benchmarking Sequential Narrative Generation in Text-to-Video Models

2025-10-14 · Zhengxu Tang, Zizheng Wang, Luning Wang, Zitao Shuai 외 arxiv

Text-to-video (T2V) generation models have made significant progress in creating visually appealing videos. However, they struggle with generating coherent sequential narratives that require logical progression through m…

Computational Efficiency

Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics

2025-09-12 · Yuriel Ryan, Rui Yang Tan, Kenny Tsu Wei Choo, Roy Ka-Wei Lee arxiv

Understanding humor is a core aspect of social intelligence, yet it remains a significant challenge for Large Multimodal Models (LMMs). We introduce PixelHumor, a benchmark dataset of 2,800 annotated multi-panel comics d…

SNaC: Coherence Error Detection for Narrative Summarization

2022-05-19 · Tanya Goyal, Junyi Jessy Li, Greg Durrett

Progress in summarizing long texts is inhibited by the lack of appropriate evaluation frameworks. When a long summary must be produced to appropriately cover the facets of that text, that summary needs to present a coher…

BenchmarkingCoherence EvaluationDocument Summarization

Fantastic Questions and Where to Find Them: FairytaleQA – An Authentic Dataset for Narrative Comprehension

2022-05-01 · ACL 2022 5 · Ying Xu, Dakuo Wang, Mo Yu, Daniel Ritchie 외

Question answering (QA) is a fundamental means to facilitate assessment and training of narrative comprehension skills for both machines and young children, yet there is scarcity of high-quality QA datasets carefully des…

BenchmarkingQuestion AnsweringQuestion GenerationQuestion-Generation