paper-with-me

홈 › Papers

Heaven-Sent or Hell-Bent? Benchmarking the Intelligence and Defectiveness of LLM Hallucinations

2025-12-25 · Chengxu Yang, Jingling Yuan, Siqi Cai, Jiawei Jiang, Chuang Hu arxiv

Hallucinations in large language models (LLMs) are commonly regarded as errors to be minimized. However, recent perspectives suggest that some hallucinations may encode creative or epistemically valuable content, a dimension that remains underquantified in current literature. Existing hallucination detection methods primarily focus on factual consistency, struggling to handle heterogeneous scientific tasks and balance creativity with accuracy. To address these challenges, we propose HIC-Bench, a novel evaluation framework that categorizes hallucinations into Intelligent Hallucinations (IH) and Defective Hallucinations (DH), enabling systematic investigation of their interplay in LLM creativity. HIC-Bench features three core characteristics: (1) Structured IH/DH Assessment. using a multi-dimensional metric matrix integrating Torrance Tests of Creative Thinking (TTCT) metrics (Originality, Feasibility, Value) with hallucination-specific dimensions (scientific plausibility, factual deviation); (2) Cross-Domain Applicability. spanning ten scientific domains with open-ended innovation tasks; and (3) Dynamic Prompt Optimization. leveraging the Dynamic Hallucination Prompt (DHP) to guide models toward creative and reliable outputs. The evaluation process employs multiple LLM judges, averaging scores to mitigate bias, with human annotators verifying IH/DH classifications. Experimental results reveal a nonlinear relationship between IH and DH, demonstrating that creativity and correctness can be jointly optimized. These insights position IH as a catalyst for creativity and reveal the ability of LLM hallucinations to drive scientific innovation.Additionally, the HIC-Bench offers a valuable platform for advancing research into the creative intelligence of LLM hallucinations.

📄 PDF Abstract BibTeX arXiv:2512.21635

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ShellfishNet: A Domain-Specific Benchmark for Visual Recognition of Marine Molluscs

2026-05-08 · Ziheng Zhou, Yang Wang, Nan Wang, Chengliang Wu 외 arxiv

The decline of global shellfish biodiversity poses a severe threat to coastal ecosystems. Although artificial intelligence (AI) technologies show potential for automated ecological monitoring, existing marine benthic dat…

Self-Supervised LearningImage Captioning

reCSE: Portable Reshaping Features for Sentence Embedding in Self-supervised Contrastive Learning

2024-08-09 · Fufangchen Zhao, Jian Gao, Danfeng Yan

We propose reCSE, a self supervised contrastive learning sentence representation framework based on feature reshaping. This framework is different from the current advanced models that use discrete data augmentation meth…

Contrastive LearningData AugmentationGPUSemantic Similarity+4

Hybrid-Vector Retrieval for Visually Rich Documents: Combining Single-Vector Efficiency and Multi-Vector Accuracy

2025-10-25 · Juyeon Kim, Geon Lee, Dongwon Choi, Taeuk Kim 외 arxiv

Retrieval over visually rich documents is essential for tasks such as legal discovery, scientific search, and enterprise knowledge management. Existing approaches fall into two paradigms: single-vector retrieval, which i…

Idiom Paraphrases: Seventh Heaven vs Cloud Nine

2015-09-01 · WS 2015 9 · Maria Pershina, Yifan He, Ralph Grishman
Natural Language InferenceParaphrase IdentificationQuestion AnsweringText Summarization

Towards detection and classification of microscopic foraminifera using transfer learning

2020-01-14 · Thomas Haugland Johansen, Steffen Aagaard Sørensen

Foraminifera are single-celled marine organisms, which may have a planktic or benthic lifestyle. During their life cycle they construct shells consisting of one or more chambers, and these shells remain as fossils in mar…

ClassificationImage ClassificationTransfer Learning