paper-with-me

홈 › Papers

Temporal Fact Conflicts in LLMs: Reproducibility Insights from Unifying DYNAMICQA and MULAN

2026-03-16 · Ritajit Dey, Iadh Ounis, Graham McDonald, Yashar Moshfeghi arxiv

Large Language Models (LLMs) often struggle with temporal fact conflicts due to outdated or evolving information in their training data. Two recent studies with accompanying datasets report opposite conclusions on whether external context can effectively resolve such conflicts. DYNAMICQA evaluates how effective external context is in shifting the model's output distribution, finding that temporal facts are more resistant to change. In contrast, MULAN examines how often external context changes memorised facts, concluding that temporal facts are easier to update. In this reproducibility paper, we first reproduce experiments from both benchmarks. We then reproduce the experiments of each study on the dataset of the other to investigate the source of their disagreement. To enable direct comparison of findings, we standardise both datasets to align with the evaluation settings of each study. Importantly, using an LLM, we synthetically generate realistic natural language contexts to replace MULAN's programmatically constructed statements when reproducing the findings of DYNAMICQA. Our analysis reveals strong dataset dependence: MULAN's findings generalise under both methodological frameworks, whereas applying MULAN's evaluation to DYNAMICQA yields mixed outcomes. Finally, while the original studies only considered 7B LLMs, we reproduce these experiments across LLMs of varying sizes, revealing how model size influences the encoding and updating of temporal facts. Our results highlight how dataset design, evaluation metrics, and model size shape LLM behaviour in the presence of temporal knowledge conflicts.

📄 PDF Abstract BibTeX arXiv:2603.15892

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Getting Sick After Seeing a Doctor? Diagnosing and Mitigating Knowledge Conflicts in Event Temporal Reasoning

2023-05-24 · Tianqing Fang, Zhaowei Wang, Wenxuan Zhou, Hongming Zhang 외

Event temporal reasoning aims at identifying the temporal relations between two or more events from narratives. However, knowledge conflicts arise when there is a mismatch between the actual temporal relations of events …

counterfactualData AugmentationHallucinationIn-Context Learning

Analysing the Residual Stream of Language Models Under Knowledge Conflicts

2024-10-21 · Yu Zhao, Xiaotang Du, Giwon Hong, Aryo Pradipta Gema 외

Large language models (LLMs) can store a significant amount of factual knowledge in their parameters. However, their parametric knowledge may conflict with the information provided in the context. Such conflicts can lead…

Reflections on the Reproducibility of Commercial LLM Performance in Empirical Software Engineering Studies

2025-10-29 · Florian Angermeir, Maximilian Amougou, Mark Kreitz, Andreas Bauer 외 arxiv

Large Language Models have gained remarkable interest in industry and academia. The increasing interest in LLMs in academia is also reflected in the number of publications on this topic over the last years. For instance,…

ConflictBank: A Benchmark for Evaluating the Influence of Knowledge Conflicts in LLM

2024-08-22 · Zhaochen Su, Jun Zhang, Xiaoye Qu, Tong Zhu 외

Large language models (LLMs) have achieved impressive advancements across numerous disciplines, yet the critical issue of knowledge conflicts, a major source of hallucinations, has rarely been studied. Only a few researc…

Misinformation

When Facts Change: Probing LLMs on Evolving Knowledge with evolveQA

2025-10-22 · Nishanth Sridhar Nakshatri, Shamik Roy, Manoj Ghuhan Arivazhagan, Hanhan Zhou 외 arxiv

LLMs often fail to handle temporal knowledge conflicts--contradictions arising when facts evolve over time within their training data. Existing studies evaluate this phenomenon through benchmarks built on structured know…