paper-with-me

홈 › Papers

Characterizing Narrative Content in Web-scale LLM Pretraining Data

2026-06-17 · Teagan Johnson, Elliott Ash, Andrew Piper, Maria Antoniak arxiv

The narrative composition of web-scale LLM pretraining corpora remains largely unexplored even though narrative is a fundamental mode of human communication. We present the first fine-grained study of narrative features in Dolma, a 3-trillion-token open pretraining corpus. Drawing on narrative theory, we design a framework spanning three core narrative elements (agency, setting, and events) operationalized as 11 interpretable dimensions. After sampling and annotating a diverse set of 400 passages, we finetune and validate NarraBERT, a RoBERTa-based model for fine-grained narrative prediction. We apply NarraBERT to 3M passages, resulting in a new dataset, NarraDolma. We find (i) narrative structure is measurable at scale across extremely heterogeneous data, (ii) we uncover a continuous, multidimensional narrative structure underlying web text, and (iii) narrative qualities are unequally distributed across pretraining sources and topics in ways that current curation practices neither measure nor account for. Our framework, dataset, and analyses provide a foundation for understanding how narrative qualities are distributed in LLM pretraining data and for studying how data composition affects narrative reasoning tasks. We publicly release NarraDolma and NarraBERT.

📄 PDF Abstract BibTeX arXiv:2606.19468

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs

2025-05-04 · Sai Krishna Mendu, Harish Yenala, Aditi Gulati, Shanu Kumar 외

Large language models (LLMs) have become integral to various real-world applications, leveraging massive, web-sourced datasets like Common Crawl, C4, and FineWeb for pretraining. While these datasets provide linguistic d…

MisinformationText Generation

OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language Pretraining

2024-11-23 · Ming Hu, Kun Yuan, Yaling Shen, Feilong Tang 외

Surgical practice involves complex visual interpretation, procedural skills, and advanced medical knowledge, making surgical vision-language pretraining (VLP) particularly challenging due to this complexity and the limit…

Representation LearningRetrieval

BIASEDTALES-ML: A Multilingual Dataset for Analyzing Narrative Attribute Distributions in LLM-Generated Stories

2026-04-18 · Yuxuan Ouyang, yingfeng luo, JingBo Zhu, Tong Xiao arxiv

Large Language Models (LLMs) are increasingly used to generate narrative content, including children's stories, which play an important role in social and cultural learning. Despite growing interest in AI safety and alig…

Characterizing Cultural Localization in AI-Generated Stories

2026-06-12 · Shaily Bhatt, Supriti Vijay, Jeremiah Milbauer, Fernando Diaz arxiv

The global use of artificial intelligence has increased interest in assessing the ability to generate culturally localized content, including stories. Cultural localization in stories often occurs through either template…

Decoding Fake Narratives in Spreading Hateful Stories: A Dual-Head RoBERTa Model with Multi-Task Learning

2025-12-18 · Yash Bhaskar, Sankalp Bahad, Parameswari Krishnamurthy arxiv

Social media platforms, while enabling global connectivity, have become hubs for the rapid spread of harmful content, including hate speech and fake narratives \cite{davidson2017automated, shu2017fake}. The Faux-Hate sha…

severity predictionMulti-Task Learning