paper-with-me

홈 › Papers

Facts Fade Fast: Evaluating Memorization of Outdated Medical Knowledge in Large Language Models

2025-09-04 · Juraj Vladika, Mahdi Dhaini, Florian Matthes arxiv

The growing capabilities of Large Language Models (LLMs) show significant potential to enhance healthcare by assisting medical researchers and physicians. However, their reliance on static training data is a major risk when medical recommendations evolve with new research and developments. When LLMs memorize outdated medical knowledge, they can provide harmful advice or fail at clinical reasoning tasks. To investigate this problem, we introduce two novel question-answering (QA) datasets derived from systematic reviews: MedRevQA (16,501 QA pairs covering general biomedical knowledge) and MedChangeQA (a subset of 512 QA pairs where medical consensus has changed over time). Our evaluation of eight prominent LLMs on the datasets reveals consistent reliance on outdated knowledge across all models. We additionally analyze the influence of obsolete pre-training data and training strategies to explain this phenomenon and propose future directions for mitigation, laying the groundwork for developing more current and reliable medical AI systems.

📄 PDF Abstract BibTeX arXiv:2509.04304

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

HALO: Half Life-Based Outdated Fact Filtering in Temporal Knowledge Graphs

2025-05-12 · Feng Ding, Tingting Wang, Yupeng Gao, Shuo Yu 외

Outdated facts in temporal knowledge graphs (TKGs) result from exceeding the expiration date of facts, which negatively impact reasoning performance on TKGs. However, existing reasoning methods primarily focus on positiv…

Knowledge Graphs

ArtiFade: Learning to Generate High-quality Subject from Blemished Images

2024-09-05 · CVPR 2025 1 · Shuya Yang, Shaozhe Hao, Yukang Cao, Kwan-Yee K. Wong

Subject-driven text-to-image generation has witnessed remarkable advancements in its ability to learn and capture characteristics of a subject using only a limited number of images. However, existing methods commonly rel…

Image GenerationText to Image GenerationText-to-Image Generation

V-DyKnow: A Dynamic Benchmark for Time-Sensitive Knowledge in Vision Language Models

2026-03-17 · Seyed Mahed Mousavi, Christian Moiola, Massimo Rizzoli, Simone Alghisi 외 arxiv

Vision-Language Models (VLMs) are trained on data snapshots of documents, including images and texts. Their training data and evaluation benchmarks are typically static, implicitly treating factual knowledge as time-inva…

knowledge editing

HoH: A Dynamic Benchmark for Evaluating the Impact of Outdated Information on Retrieval-Augmented Generation

2025-03-03 · Jie Ouyang, Tingyue Pan, Mingyue Cheng, Ruiran Yan 외

While Retrieval-Augmented Generation (RAG) has emerged as an effective approach for addressing the knowledge outdating problem in Large Language Models (LLMs), it faces a critical challenge: the prevalence of outdated in…

RAGRetrievalRetrieval-augmented Generation

Deep Outdated Fact Detection in Knowledge Graphs

2024-02-06 · Huiling Tu, Shuo Yu, Vidya Saikrishna, Feng Xia 외

Knowledge graphs (KGs) have garnered significant attention for their vast potential across diverse domains. However, the issue of outdated facts poses a challenge to KGs, affecting their overall quality as real-world inf…

Knowledge Graphs