paper-with-me

Papers

Quality Without Usefulness: LLM-Generated XAI Narratives as Trust Heuristics Rather Than Decision Aids

2026-05-26 · Fabian Lukassen, Jan Herrmann, Christoph Weisser, Alexander Silbersdorff, Benjamin Saefken, Thomas Kneib arxiv

Prior work shows that Large Language Models (LLMs) can transform Explainable AI (XAI) outputs into Natural Language Explanations (NLEs) that score highly on quality metrics such as plausibility, coherence, and comprehensibility. But does explanation quality translate to practical usefulness? We investigate this question in a time-series energy forecasting domain through five controlled experiments (2,730 judgments across 60 test instances), each operationalising a distinct facet of usefulness studied in the XAI literature. Holding NLE quality constant at the high levels established by a prior factorial study, we find that NLEs do not improve task accuracy on any of the five tasks, while inflating self-reported confidence. A placebic control shows that this confidence boost is driven by text presence rather than content. In an out-of-distribution detection task, NLEs reduce the LLM judge's ability to flag unreliable predictions, providing false reassurance that masks model failure. We characterise these findings as the Quality-Usefulness Gap and argue that evaluation of the XAI-to-NLE pipeline must extend beyond text-quality metrics to downstream task performance.

📄 PDF Abstract BibTeX arXiv:2605.26770

Code (0)

등록된 구현이 없습니다.

Tasks

Out-of-Distribution Detection

Similar Papers 제목 키워드 기반

A Novel Corpus of Discourse Structure in Humans and Computers

2021-11-10 · Babak Hemmatian, Sheridan Feucht, Rachel Avram, Alexander Wey 외

We present a novel corpus of 445 human- and computer-generated documents, comprising about 27,000 clauses, annotated for semantic clause types and coherence relations that allow for nuanced comparison of artificial and n…

Text Generation

When Readability and Source Retention Diverge: An Evaluability Gap in AI Translation

2026-08-19 · Chenchen Mao, Hanjing Shi, Haiyan Jia, Emily Wegrzyn 외 arxiv

Readable AI output can leave an evaluability gap: even when the source is shown, an overall-quality judgment may not reflect what an output preserves. We investigated how source-text condition and output rendering relate…

A Reflective Storytelling Agent for Older Adults: Integrating Argumentation Schemes and Argument Mining in LLM-Based Personalised Narratives

2026-05-11 · Jayalakshmi Baskar, Vera C. Kaelin, Kaan Kilic, Helena Lindgren arxiv

This work investigates whether knowledge-driven large language model (LLM)-based storytelling can support purposeful narrative interaction with a digital companion for older adults. To address known limitations of LLMs, …

Knowledge GraphsArgument Mining

Evaluating Quality of Gaming Narratives Co-created with AI

2025-09-04 · Arturo Valdivia, Paolo Burelli arxiv

This paper proposes a structured methodology to evaluate AI-generated game narratives, leveraging the Delphi study structure with a panel of narrative design experts. Our approach synthesizes story quality dimensions fro…

Human-AI Narrative Synthesis to Foster Shared Understanding in Civic Decision-Making

2025-09-23 · Cassandra Overney, Hang Jiang, Urooj Haider, Cassandra Moe 외 arxiv

Community engagement processes in representative political contexts, like school districts, generate massive volumes of feedback that overwhelm traditional synthesis methods, creating barriers to shared understanding not…