paper-with-me

홈 › Papers

Precision Is Not Faithfulness: Coverage-Aware Evaluation of Grounded Generation with a Complete Oracle

2026-06-08 · Juan S. Santillana arxiv

Reference-free faithfulness metrics verify each atomic claim a model makes against ground truth, and are increasingly used to evaluate grounded generation. We show they share a blind spot: they measure only precision -- are the stated claims supported? -- and therefore reward abstention, since a model can score near-perfect faithfulness by saying almost nothing. We make this measurable using Formula 1 telemetry, a domain where strategic ground truth is derived deterministically and, crucially, completely: for each decision we know the full set of facts that mattered. This completeness -- absent in open-domain faithfulness benchmarks -- lets us measure recall (coverage of the relevant facts) exactly, alongside precision. On a multilingual (EN/ES/PT) benchmark of 7,253 decision instances spanning 157 races, the most precise frontier model covers under half of the relevant facts and ranks last by F1, so requiring coverage reorders the systems; the same effect reappears in a second complete-oracle domain (NOAA weather forecasts). Fine-tuning small models (1B-7B) on the complete oracle closes the precision-recall gap entirely (F1 ~0.98), beating every zero-shot frontier system regardless of scale. We pair faithfulness with coverage into a single score, validate the metric (controlled perturbation; agreement across a model-free regex extractor and a cross-family LLM extractor, system-level Spearman 1.0), and give a verifier-guided generation method that improves precision and recall without references. We release the benchmark, structured annotations, metric, baselines, and an interactive demo.

📄 PDF Abstract BibTeX arXiv:2606.09376

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

KG-First, LLM-Fallback: A Hybrid Microservice for Grounded Skill Search and Explanation

2026-05-02 · Ngoc Luyen Le, Marie-Hélène Abel, Bertrand Laforge arxiv

Authoritative competency frameworks such as ESCO, ROME, and O*NET are essential for aligning education with labor market needs, yet their technical complexity and structural heterogeneity hinder practical adoption by edu…

TRIAGE: Trustworthy Retrieval Instrumentation And Graph Evaluation

2026-07-03 · Axel TahmasebiMoradi, Lucas Schott, Martin Royer arxiv

Knowledge graphs (KGs) that underpin Graph-based Retrieval-Augmented Generation (Graph-RAG) are increasingly built automatically by LLM-driven extraction rather than curated by experts. Proper evaluation would require in…

Knowledge Graphs

FFCI: A Framework for Interpretable Automatic Evaluation of Summarization

2020-11-27 · Fajri Koto, Timothy Baldwin, Jey Han Lau

In this paper, we propose FFCI, a framework for fine-grained summarization evaluation that comprises four elements: faithfulness (degree of factual consistency with the source), focus (precision of summary content relati…

Question AnsweringSemantic Textual SimilaritySentenceSTS

Toward Faithful and Complete Answer Construction from a Single Document

2026-02-05 · Zhaoyang Chen, Cody Fleming arxiv

Modern large language models (LLMs) are powerful generators driven by statistical next-token prediction. While effective at producing fluent text, this design biases models toward high-probability continuations rather th…

DrugClaw and DrugAudit: A Primary-Source-Grounded Agent and Authority-Aware Benchmark for Drug-Information Question Answering

2026-05-31 · Qing Wang, Bo Li, Jialu Liang, Daling Shi 외 arxiv

Drug-information question answering is a high-stakes setting where hallucinated facts can mislead clinical decision-making and the provenance of each cited fact matters as much as the fact itself. We present DrugClaw, a …

Question Answering