paper-with-me

홈 › Papers

MedFabric: Gold Evidence Hides the Difficulty of Word-Level Medical Fabrication Detection

2026-05-05 · Tung Sum Thomas Kwok, Qian Qian, Xiaofeng Lin, Dongxu Zhang, Jun Han, Zhichao Yang, Davin Hill, Tamer Soliman, Sanjit Singh Batra, Robert Tillman, Guang Cheng arxiv

Large language models fabricate in medicine, producing fluent statements that are factually wrong, so reliable fabrication detection is a prerequisite for clinical deployment. Reported progress on this task is inflated by two evaluation artifacts: an authorship-style shortcut, where human-written ground truths are paired with LLM-written hallucinations so detectors key on writing style rather than facts, and the provision of gold evidence at test time. A benchmark that tests factual reasoning must therefore remove the style shortcut, ground every fabrication in a real retrievable passage, and score detectors across the range of evidence quality faced in deployment. We build MedFabric to these requirements, a benchmark of 646 word-level medical fabrications, each paired with a ground truth that shares its LLM authorship and near-identical surface form (median ROUGE-L 0.95). On MedFabric the task is unsolved: expert clinicians reach only 53.3% macro F1 and no detector family clears about 60% without gold evidence. Our central finding is that detection is governed by evidence correctness rather than fabrication subtlety, since a strong LLM scores 91% with the gold passage but falls to 35%, below its own no-evidence baseline, under a wrong one, a pattern that holds on two benchmarks at two model scales. The failure is actionable: a retrieval-confidence gate that abstains on low-confidence evidence raises macro F1 from 61% to 74%, and we release MedFabric, all code, and every baseline.

📄 PDF Abstract BibTeX arXiv:2605.04180

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

EviScope: Paired Counterfactual Evidence Diagnostics for Faithful and Efficient Grounded Language Models

2026-09-15 · Suryadeep Singh Deswal arxiv

Grounded language-model systems are often evaluated by final answer accuracy, yet a correct answer can be unsupported, drawn from the wrong source, or produced when evidence is insufficient or contradictory. We introduce…

CEFR-Annotated WordNet: LLM-Based Proficiency-Guided Semantic Database for Language Learning

2025-10-21 · Masato Kikuchi, Masatsugu Ono, Toshioki Soga, Tetsu Tanabe 외 arxiv

Although WordNet is a valuable resource because of its structured semantic networks and extensive vocabulary, its fine-grained sense distinctions can be challenging for second-language learners. To address this issue, we…

Semantic Similarity

Using semi-experts to derive judgments on word sense alignment: a pilot study

2012-05-01 · LREC 2012 5 · Soojeong Eom, Markus Dickinson, Graham Katz

The overall goal of this project is to evaluate the performance of word sense alignment (WSA) systems, focusing on obtaining examples appropriate to language learners. Building a gold standard dataset based on human expe…

Machine TranslationWord Sense Disambiguation

Evidence-Grounded Ensemble Diagnosis of 802.11 Packet Captures: A Multi-Stage Pipeline with Deterministic Reliability Scoring

2026-06-05 · Jerome Henry, Swadhin Pradhan, Miroslav Popovic arxiv

Diagnosing 802.11 packet captures requires expert protocol knowledge, is slow, inconsistent across engineers, and unscalable. LLM-based approaches sound plausible but fabricate protocol events absent from captures (espec…

Prior over Evidence: Stereotype-Driven Diagnosis in LLM-Based L2 Pronunciation Feedback

2026-06-13 · Rong Wang, Kun Sun arxiv

Large language models are increasingly deployed for written pronunciation feedback in second-language (L2) English learning, under the assumption that their diagnoses are grounded in the supplied speech evidence rather t…