paper-with-me

홈 › Papers

Critical Confabulation: Can LLMs Hallucinate for Social Good?

2025-11-11 · Peiqi Sui, Eamon Duede, Hoyt Long, Richard Jean So arxiv

LLMs hallucinate, yet some confabulations can have social affordances if carefully bounded. We propose critical confabulation (inspired by critical fabulation from literary and social theory), the use of LLM hallucinations to "fill-in-the-gap" for omissions in archives due to social and political inequality, and reconstruct divergent yet evidence-bound narratives for history's ``hidden figures''. We simulate these gaps with an open-ended narrative cloze task: asking LLMs to generate a masked event in a character-centric timeline sourced from a novel corpus of unpublished texts. We evaluate audited (for data contamination), fully-open models (the OLMo-2 family) and unaudited open-weight and proprietary baselines under a range of prompts designed to elicit controlled and useful hallucinations. Our findings validate LLMs' foundational narrative understanding capabilities to perform critical confabulation, and show how controlled and well-specified hallucinations can support LLM applications for knowledge production without collapsing speculation into a lack of historical accuracy and fidelity.

📄 PDF Abstract BibTeX arXiv:2511.07722

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Confabulation: The Surprising Value of Large Language Model Hallucinations

2024-06-06 · Peiqi Sui, Eamon Duede, Sophie Wu, Richard Jean So

This paper presents a systematic defense of large language model (LLM) hallucinations or 'confabulations' as a potential resource instead of a categorically negative pitfall. The standard view is that confabulations are …

HallucinationLanguage ModelingLanguage ModellingLarge Language Model+1

Detecting hallucinations in large language models using semantic entropy

2024-06-19 · Nature 2024 6 · Sebastian Farquhar, Jannik Kossen, Lorenz Kuhn, Yarin Gal

Large language model (LLM) systems, such as ChatGPT1 or Gemini2, can show impressive reasoning and question-answering capabilities but often ‘hallucinate’ false outputs and unsubstantiated answers3,4. Answering unreliabl…

Large Language ModelQuestion Answering

Memory Helps, but Confabulation Misleads: Understanding Streaming Events in Videos with MLLMs

2025-02-21 · Gengyuan Zhang, Mingcong Ding, Tong Liu, Yao Zhang 외

Multimodal large language models (MLLMs) have demonstrated strong performance in understanding videos holistically, yet their ability to process streaming videos-videos are treated as a sequence of visual events-remains …

Misinformation

ReFACT: A Benchmark for Scientific Confabulation Detection with Positional Error Annotations

2025-09-30 · Yindong Wang, Martin Preiß, Margarita Bugueño, Jan Vincent Hoffbauer 외 arxiv

The mechanisms underlying scientific confabulation in Large Language Models (LLMs) remain poorly understood. We introduce ReFACT (Reddit False And Correct Texts), a benchmark of 1,001 expert-annotated question-answer pai…

Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality?

2026-05-21 · Caixin Kang, Tianyu Yan, Sitong Gong, Mingfang Zhang 외 arxiv

Multimodal Large Language Models (MLLMs) are increasingly deployed in human-facing roles where personality perception is critical, yet existing benchmarks evaluate this capability solely on numerical Big Five score predi…