paper-with-me

홈 › Papers

The Future of Facts: Tracing the Factual Generation-Verification Gap

2026-05-26 · Tim R. Davidson, Anja Surina, Caglar Gulcehre arxiv

Language models are becoming the default interface to factual knowledge, yet they often verify outputs more reliably than they generate them. This generation-verification gap (GV-gap) underlies many recent advances in self-improvement and reasoning, but its dynamics on factual knowledge specifically remain poorly understood. We focus on the training mechanisms underlying factual GV-gaps, distinguishing them from their computational and aesthetic counterparts. We trace generation and verification capabilities through three training phases (acquisition, continual learning, and updating) across four open-source model families at two scales each. Three findings recur across models: (i) verification is consistently learned before generation; (ii) verification is more robust to continual learning than generation; and (iii) factual updates can leave models in a "multi-verse" state, simultaneously verifying both old and new answers as correct. Natural experiments on frontier models reproduce these dynamics at scale and reveal residual verification biases on well-covered facts.

📄 PDF Abstract BibTeX arXiv:2605.27564

Code (0)

등록된 구현이 없습니다.

Tasks

Continual Learning

Similar Papers 제목 키워드 기반

Not All Claims Are Equally Risky: FACTOR for Adaptive Verification in Factual Long-Form Generation

2026-06-21 · Areeba Hassan, Arooj Kausar, Syeda Kisaa Fatima, Gibrail Islam 외 arxiv

Large Language Models (LLMs) generate fluent long-form text, however, often add unsupported factual claims. Existing verification techniques improve factuality by grounding generation in external evidence. However, the s…

DnDScore: Decontextualization and Decomposition for Factuality Verification in Long-Form Text Generation

2024-12-17 · Miriam Wanner, Benjamin Van Durme, Mark Dredze

The decompose-then-verify strategy for verification of Large Language Model (LLM) generations decomposes claims that are then independently verified. Decontextualization augments text (claims) to ensure it can be verifie…

FormLanguage ModelingLanguage ModellingLarge Language Model+1

Expert-Aware Causal Tracing of Factual Recall in Sparse MoE Language Models

2026-06-02 · Yuetian Lu, Ali Modarressi, Yihong Liu, Hinrich Schütze arxiv

Causal tracing of factual recall has been studied predominantly in dense transformer language models, where interventions localize information flow to layers or feed-forward modules. Sparse mixture-of-experts (MoE) langu…

Mitigating LLM Hallucinations through Domain-Grounded Tiered Retrieval

2026-03-18 · Md. Asraful Haque, Aasar Mehdi, Maaz Mahboob, Tamkeen Fatima arxiv

Large Language Models (LLMs) have achieved unprecedented fluency but remain susceptible to "hallucinations" - the generation of factually incorrect or ungrounded content. This limitation is particularly critical in high-…

Global Facts

MedScore: Factuality Evaluation of Free-Form Medical Answers

2025-05-24 · Heyuan Huang, Alexandra DeLucia, Vijay Murari Tiyyala, Mark Dredze

While Large Language Models (LLMs) can generate fluent and convincing responses, they are not necessarily correct. This is especially apparent in the popular decompose-then-verify factuality evaluation pipeline, where LL…

FormHallucinationSentencevalid