paper-with-me

홈 › Papers

Hallucination or Creativity: How to Evaluate AI-Generated Scientific Stories?

2026-02-02 · Alex Argese, Pasquale Lisena, Raphaël Troncy arxiv

Generative AI can turn scientific articles into narratives for diverse audiences, but evaluating these stories remains challenging. Storytelling demands abstraction, simplification, and pedagogical creativity-qualities that are not often well-captured by standard summarization metrics. Meanwhile, factual hallucinations are critical in scientific contexts, yet, detectors often misclassify legitimate narrative reformulations or prove unstable when creativity is involved. In this work, we propose StoryScore, a composite metric for evaluating AI-generated scientific stories. StoryScore integrates semantic alignment, lexical grounding, narrative control, structural fidelity, redundancy avoidance, and entity-level hallucination detection into a unified framework. Our analysis also reveals why many hallucination detection methods fail to distinguish pedagogical creativity from factual errors, highlighting a key limitation: while automatic metrics can effectively assess semantic similarity with original content, they struggle to evaluate how it is narrated and controlled.

📄 PDF Abstract BibTeX arXiv:2602.02290

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic Similarity

Similar Papers 제목 키워드 기반

Evaluating Creative Short Story Generation in Humans and Large Language Models

2024-11-04 · Mete Ismayilzada, Claire Stevenson, Lonneke van der Plas

Story-writing is a fundamental aspect of human imagination, relying heavily on creativity to produce narratives that are novel, effective, and surprising. While large language models (LLMs) have demonstrated the ability …

DiversitySentenceStory Generation

Art or Artifice? Large Language Models and the False Promise of Creativity

2023-09-25 · Tuhin Chakrabarty, Philippe Laban, Divyansh Agarwal, Smaranda Muresan 외

Researchers have argued that large language models (LLMs) exhibit high-quality writing capabilities from blogs to stories. However, evaluating objectively the creativity of a piece of writing is challenging. Inspired by …

Forma mentis networks predict creativity ratings of short texts via interpretable artificial intelligence in human and GPT-simulated raters

2024-11-30 · Edith Haim, Natalie Fischer, Salvatore Citraro, Giulio Rossetti 외

Creativity is a fundamental skill of human cognition. We use textual forma mentis networks (TFMN) to extract network (semantic/syntactic associations) and emotional features from approximately one thousand human- and GPT…

Explainable artificial intelligenceExplainable Artificial Intelligence (XAI)Feature Importance

Heaven-Sent or Hell-Bent? Benchmarking the Intelligence and Defectiveness of LLM Hallucinations

2025-12-25 · Chengxu Yang, Jingling Yuan, Siqi Cai, Jiawei Jiang 외 arxiv

Hallucinations in large language models (LLMs) are commonly regarded as errors to be minimized. However, recent perspectives suggest that some hallucinations may encode creative or epistemically valuable content, a dimen…

Does Less Hallucination Mean Less Creativity? An Empirical Investigation in LLMs

2025-12-12 · Mohor Banerjee, Nadya Yuki Wangsajaya, Syed Ali Redha Alsagoff, Min Sen Tan 외 arxiv

Large Language Models (LLMs) exhibit remarkable capabilities in natural language understanding and reasoning, but suffer from hallucination: the generation of factually incorrect content. While numerous methods have been…

Natural Language Understanding