paper-with-me

홈 › Papers

NarrativeWorldBench: A Frontier-Saturated Benchmark and a Latent World Model for Long-Horizon Co-Creative Audio Drama

2026-06-16 · Logan Mann, Abdur Rahman, Mohammad Saifullah, Taaha Kazi, Vasu Sharma arxiv

Long-form serialized audio drama, with arcs that run for 200 to 800 episodes, is a major creative medium and a setting where frontier large language models (LLMs) fail. We benchmark 21 models, spanning classical, fine-tuned, open-frontier, closed-frontier, and reasoning tiers, on a uniform set of structural narrative metrics. All closed-frontier systems saturate at a plot-beat F1 in the band [0.78, 0.81] and collapse by about -0.20 F1 at horizon h=200. We introduce NarrativeWorldBench, an open benchmark of nine narrative-structure metrics evaluated across horizons h in {10, 20, 50, 100, 200}, with cross-lingual evaluation across four Indic languages (Hindi, Tamil, Telugu, Marathi). We introduce N-VSSM, a Narrative Variational State-Space Model that maintains a structured 256-dimensional latent world state over more than 200 episodes via a Mamba-2 backbone with an event-conditioned posterior and an 8B decoder. N-VSSM holds plot-beat F1 >= 0.84 across all horizons at 4x lower compute than the closed-frontier band. A learned Cultural Transfer Function lifts cross-language fidelity by +0.20 to +0.23 Likert points. In a within-subjects writer study (n = 12 professional authors, 240 trials), N-VSSM is preferred over Claude Opus 4.5 on long-arc consistency 71% of the time and rated +1.3 Likert points higher on controllability.

📄 PDF Abstract BibTeX arXiv:2606.17391

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SEAL: Can Saturated Benchmarks Be Revived by LLM-as-a-Meta-Judge?

2026-05-28 · Jiamin Chen, Yidi Wu, Qiexiang Wang, Qianben Chen 외 arxiv

Widely used language-model benchmarks are increasingly saturated, with frontier systems often receiving near-tied scores that standard metrics cannot resolve. Rather than constructing harder alternatives, we ask whether …

Mathematical ReasoningQuestion AnsweringCode Generation

Deep Richardson-Lucy Deconvolution for Low-Light Image Deblurring

2023-08-10 · Liang Chen, Jiawei Zhang, Zhenhua Li, Yunxuan Wei 외

Images taken under the low-light condition often contain blur and saturated pixels at the same time. Deblurring images with saturated pixels is quite challenging. Because of the limited dynamic range, the saturated pixel…

DeblurringImage Deblurring

GeMA: Learning Latent Manifold Frontiers for Benchmarking Complex Systems

2026-03-17 · Jia Ming Li, Anupriya, Daniel J. Graham arxiv

Benchmarking the performance of complex systems such as rail networks, renewable generation assets and national economies is central to transport planning, regulation and macroeconomic analysis. Classical frontier method…

Blind Deblurring for Saturated Images

2021-06-19 · CVPR 2021 1 · Liang Chen, Jiawei Zhang, Songnan Lin, Faming Fang 외

Blind deblurring has received considerable attention in recent years. However, state-of-the-art methods often fail to process saturated blurry images. The main reason is that saturated pixels are not conforming to th…

Deblurring

FrontierScience: Evaluating AI's Ability to Perform Expert-Level Scientific Tasks

2026-01-29 · Miles Wang, Robi Lin, Kat Hu, Joy Jiao 외 arxiv

We introduce FrontierScience, a benchmark evaluating expert-level scientific reasoning in frontier language models. Recent model progress has nearly saturated existing science benchmarks, which often rely on multiple-cho…