paper-with-me

홈 › Papers

Hallucination vs interpretation: rethinking accuracy and precision in AI-assisted data extraction for knowledge synthesis

2025-08-13 · Xi Long, Christy Boscardin, Lauren A. Maggio, Joseph A. Costello, Ralph Gonzales, Rasmyah Hammoudeh, Ki Lai, Yoon Soo Park, Brian C. Gin arxiv

Knowledge syntheses (literature reviews) are essential to health professions education (HPE), consolidating findings to advance theory and practice. However, they are labor-intensive, especially during data extraction. Artificial Intelligence (AI)-assisted extraction promises efficiency but raises concerns about accuracy, making it critical to distinguish AI 'hallucinations' (fabricated content) from legitimate interpretive differences. We developed an extraction platform using large language models (LLMs) to automate data extraction and compared AI to human responses across 187 publications and 17 extraction questions from a published scoping review. AI-human, human-human, and AI-AI consistencies were measured using interrater reliability (categorical) and thematic similarity ratings (open-ended). Errors were identified by comparing extracted responses to source publications. AI was highly consistent with humans for concrete, explicitly stated questions (e.g., title, aims) and lower for questions requiring subjective interpretation or absent in text (e.g., Kirkpatrick's outcomes, study rationale). Human-human consistency was not higher than AI-human and showed the same question-dependent variability. Discordant AI-human responses (769/3179 = 24.2%) were mostly due to interpretive differences (18.3%); AI inaccuracies were rare (1.51%), while humans were nearly three times more likely to state inaccuracies (4.37%). Findings suggest AI variability depends more on interpretability than hallucination. Repeating AI extraction can identify interpretive complexity or ambiguity, refining processes before human review. AI can be a transparent, trustworthy partner in knowledge synthesis, though caution is needed to preserve critical human insights.

📄 PDF Abstract BibTeX arXiv:2508.09458

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Beyond Accuracy: Rethinking Hallucination and Regulatory Response in Generative AI

2025-09-12 · Zihao Li, Weiwei Yi, Jiahong Chen arxiv

Hallucination in generative AI is often treated as a technical failure to produce factually correct output. Yet this framing underrepresents the broader significance of hallucinated content in language models, which may …

LCAi: Life Cycle Assessment with big data fusion and retrieval-augmented generation-assisted interpretation

2026-06-25 · Georgios Tsironis, Juan D. Medrano-Garcia, Gonzalo Guillen-Gosalbez arxiv

The interpretation phase of life cycle assessment often lacks structured mechanisms for translating quantified improvement opportunities addressing environmental hotspots into actionable strategic pathways under technolo…

Fast and accurate classification of echocardiograms using deep learning

2017-06-27 · Ali Madani, Ramy Arnaout, Mohammad Mofrad, Rima Arnaout

Echocardiography is essential to modern cardiology. However, human interpretation limits high throughput analysis, limiting echocardiography from reaching its full clinical and research potential for precision medicine. …

ClassificationDeep LearningGeneral ClassificationOverall - Test

Hallucination Behavior in Multimodal LLMs Across Agricultural Image Interpretation and Generation Tasks

2026-05-26 · Partho Ghose, Al Bashir, Prem Raj, Azlan Zahid arxiv

Large Language Models (LLMs) are being rapidly adopted in agricultural imaging applications, ranging from crop interpretation to synthetic field image generation. However, these models frequently exhibit hallucinations o…

Visual ReasoningImage Generation

Rethinking Autonomy: Preventing Failures in AI-Driven Software Engineering

2025-08-15 · Satyam Kumar Navneet, Joydeep Chandra arxiv

The integration of Large Language Models (LLMs) into software engineering has revolutionized code generation, enabling unprecedented productivity through promptware and autonomous AI agents. However, this transformation …

Prompt EngineeringCode Generation