paper-with-me

홈 › Papers

SAGE: Hierarchical LLM-Based Literary Evaluation through Ontology-Grounded Interpretive Dimensions

2026-05-08 · Tianyu Wang, Nianjun Zhou arxiv

Evaluating literary quality requires assessing interpretive dimensions such as cultural representation, emotional depth, and philosophical sophistication that resist straightforward computational measurement. We introduce SAGE, a hierarchical evaluation framework that decomposes literary quality into ontology-grounded interpretive dimensions assessed through structured large language model evaluation with multi-round iterative reflection and independent validation. We validate the framework on 100 short stories (50 canonical works, 30 pulp fiction, 20 LLM-generated narratives) across three analytical layers (cultural, emotional-psychological, existential-philosophical) using dual-mode assessment. Across 600 evaluations, the framework achieves 98.8% score convergence and greater than 94% inter-rater agreement, with near-perfect mode invariance between content-based and metadata-based evaluation. Statistical analysis reveals a consistent genre hierarchy (Canonical > Pulp > LLM, all p<0.001) with layer-specific discrimination: cultural critique and philosophical depth exhibit very large effect sizes (Cohen's d>2.4), while emotional representation shows smaller gaps (d=1.68), suggesting that affective patterns are more learnable from training data than critical stance or philosophical depth. Cross-layer correlations (r=0.649-0.683) confirm the three dimensions capture empirically distinguishable quality facets. These findings demonstrate that theory-driven LLM evaluation can achieve measurement-grade reliability and support systematic identification of where current generative models fall short of human literary production, with direct implications for scalable automated evaluation of open-ended text generation.

📄 PDF Abstract BibTeX arXiv:2605.07102

Code (0)

등록된 구현이 없습니다.

Tasks

Text Generation

Similar Papers 제목 키워드 기반

The Literary Theme Ontology for Media Annotation and Information Retrieval

2019-05-01 · Paul Sheridan, Mikael Onsjö, Janna Hastings

Literary theme identification and interpretation is a focal point of literary studies scholarship. Classical forms of literary scholarship, such as close reading, have flourished with scarcely any need for commonly defin…

Information RetrievalRetrieval

An Ontology-Based Recommender System with an Application to the Star Trek Television Franchise

2018-07-31 · Paul Sheridan, Mikael Onsjö, Claudia Becerra, Sergio Jimenez 외

Collaborative filtering based recommender systems have proven to be extremely successful in settings where user preference data on items is abundant. However, collaborative filtering algorithms are hindered by their weak…

Collaborative FilteringRecommendation SystemsTopic Models

Lotte and Annette: A Framework for Finding and Exploring Key Passages in Literary Works

2021-12-01 · NLP4DH (ICON) 2021 12 · Frederik Arnold, Robert Jäschke

We present an approach that leverages expert knowledge contained in scholarly works to automatically identify key passages in literary works. Specifically, we extend a text reuse detection method for finding quotations, …

RELiC: Retrieving Evidence for Literary Claims

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Humanities scholars commonly provide evidence for claims that they make about a work of literature (e.g., a novel) in the form of quotations from the work. We collect a large-scale dataset (RELiC) of 90K literary quotati…

Information RetrievalRetrievalSemantic SimilaritySemantic Textual Similarity

Retrieval-Augmented Guardrails for AI-Drafted Patient-Portal Messages: Error Taxonomy Construction and Large-Scale Evaluation

2025-09-26 · Wenyuan Chen, Fateme Nateghi Haredasht, Kameron C. Black, Francois Grolleau 외 arxiv

Asynchronous patient-clinician messaging via EHR portals is a growing source of clinician workload, prompting interest in large language models (LLMs) to assist with draft responses. However, LLM outputs may contain clin…