paper-with-me

홈 › Papers

Errors in AI-Assisted Retrieval of Medical Literature: A Comparative Study

2026-03-21 · Jenny Gao, Yongfeng Zhang, Mary L Disis, Lanjing Zhang arxiv

Large language models (LLMs) assisted literature retrieval may lead to erroneous references, but these errors have not been rigorously quantified. Therefore, we quantitatively assess errors in reference retrieval of widely used free-version LLM platforms and identify the factors associated with retrieval errors. We evaluated 2,000 references retrieved by 5 LLMs (Grok-2, ChatGPT GPT-4.1, Google Gemini Flash 2.5, Perplexity AI, and DeepSeek GPT-4) for 40 randomly-selected original articles (10 per journal) published Jan. 2024 to July 2025 from British Medical Journal (BMJ), Journal of the American Medical Association, and The New England Journal of Medicine (NEJM). Primary outcomes were a multimetric score ratio combining validity of digital object identifier, PubMed ID, Google-Scholar link, and relevance; and complete miss rate (proportion of references failing all applicable metrics). Multivariable regression was used to examine independent associations. LLM platforms completely failed to retrieve correct reference data 47.8% of the time. The average score ratio of the 5 LLM platforms was 0.29 (standard deviation, 0.35; range, 0-1.25), with a higher score ratio indicating a higher accuracy in retrieving relevant references and correct bibliographic data. The highest and lowest accuracies were achieved by Grok (0.57) and Genimi (0.11), respectively. Compared with BMJ, NEJM articles had lower score ratios and higher complete miss rates. Multivariable analysis shows LLM platforms and journals were independently associated with score ratios and complete miss rate, respectively. We show modest overall performance of LLMs and significant variability in retrieval accuracy across platforms and journals. LLM platforms and journals are associated with LLM's performance in retrieving medical literature. Bibliographic data should be carefully reviewed when using LLM-assisted literature retrieval.

📄 PDF Abstract BibTeX arXiv:2603.22344

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Comparative Analysis of Errors in MT Output and Computer-assisted Translation: Effect of the Human Factor

2019-08-01 · WS 2019 8 · Irina Ovchinnikova, Daria Morozova
Translation

Improving Reliability and Explainability of Medical Question Answering through Atomic Fact Checking in Retrieval-Augmented LLMs

2025-05-30 · Juraj Vladika, Annika Domres, Mai Nguyen, Rebecca Moser 외

Large language models (LLMs) exhibit extensive medical knowledge but are prone to hallucinations and inaccurate citations, which pose a challenge to their clinical adoption and regulatory compliance. Current methods, suc…

Fact CheckingHallucinationLong Form Question AnsweringMedical Question Answering+2

Computational-Assisted Systematic Review and Meta-Analysis (CASMA): Effect of a Subclass of GnRH-a on Endometriosis Recurrence

2025-09-20 · Sandro Tsang arxiv

Background: Evidence synthesis facilitates evidence-based medicine. This task becomes increasingly difficult to accomplished with applying computational solutions, since the medical literature grows at astonishing rates.…

Information Retrieval

Artificial Intelligence Model for Tumoral Clinical Decision Support Systems

2023-01-09 · Guillermo Iglesias, Edgar Talavera, Jesús Troya Garcìa, Alberto Díaz-Álvarez 외

Comparative diagnostic in brain tumor evaluation makes possible to use the available information of a medical center to compare similar cases when a new patient is evaluated. By leveraging Artificial Intelligence models,…

Decision MakingDiagnosticImage RetrievalMedical Diagnosis+1

M-Eval: A Heterogeneity-Based Framework for Multi-evidence Validation in Medical RAG Systems

2025-10-28 · Mengzhou Sun, Sendong Zhao, Jianyu Chen, Haochun Wang 외 arxiv

Retrieval-augmented Generation (RAG) has demonstrated potential in enhancing medical question-answering systems through the integration of large language models (LLMs) with external medical literature. LLMs can retrieve …