paper-with-me

홈 › Papers

Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth

2026-05-24 · Yoav Gur-Arieh, Ana Marasović, Mor Geva arxiv

Chains of thought (CoTs) have become central in interpreting and auditing behaviors of large language models. Yet growing evidence suggests that these traces often fail to faithfully represent the computations behind a model's predictions. Several faithfulness metrics have been proposed, but whether they indeed measure faithfulness remains unknown. Answering this requires ground-truth labels, which are hard to obtain since internal computations are not directly observable. Consequently, most works proposing metrics report only absolute scores or comparisons to prior metrics, and the few existing benchmarks rely on proxies like plausibility or importance, properties orthogonal to faithfulness that can mislead about whether a CoT can be trusted. We address this challenge by constructing tasks whose outputs reveal which intermediate computations must have produced them, and developing an automated labeling pipeline that yields ground-truth faithfulness labels at both the step and CoT level. Building on this methodology, we present BonaFide, a benchmark of 3,066 labeled CoTs across 13 tasks and 10 models, and use it to conduct the first systematic evaluation of prominent faithfulness metrics. Our experiments show that most metrics perform near chance, exhibit strong prediction biases and degrade on longer CoTs. The best metric reaches only 0.70 AUROC at the CoT level while another reaches 0.59 at the step level, with neither transferring across settings, while entailing prohibitively high computational cost. Our results expose fundamental gaps in current faithfulness evaluation and call for the development of more reliable and efficient metrics.

📄 PDF Abstract BibTeX arXiv:2605.25052

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

BUMP: A Benchmark of Unfaithful Minimal Pairs for Meta-Evaluation of Faithfulness Metrics

2022-12-20 · Liang Ma, Shuyang Cao, Robert L. Logan IV, Di Lu 외

The proliferation of automatic faithfulness metrics for summarization has produced a need for benchmarks to evaluate them. While existing benchmarks measure the correlation with human judgements of faithfulness on model-…

A Comparative Study of Faithfulness Metrics for Model Interpretability Methods

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Interpretable methods to reveal the internal reasoning processes behind machine learning models have attracted increasing attention in recent years. To quantify the extent to which the identified interpretations truly re…

Decision Making

Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimization

2026-05-24 · Jingyi Sun, Qianli Wang, Pepa Atanasova, Nils Feldhus 외 arxiv

Chain-of-Thought (CoT) faithfulness, i.e., whether CoTs genuinely reflect large language models' (LLM) underlying behavior, is typically evaluated with metrics under two disjoint paradigms: contextual faithfulness, measu…

A Comparative Study of Faithfulness Metrics for Model Interpretability Methods

2022-04-12 · ACL 2022 5 · Chun Sik Chan, Huanqi Kong, Guanqing Liang

Interpretation methods to reveal the internal reasoning processes behind machine learning models have attracted increasing attention in recent years. To quantify the extent to which the identified interpretations truly r…

Decision Making

A Meta-Evaluation of Faithfulness Metrics for Long-Form Hospital-Course Summarization

2023-03-07 · Griffin Adams, Jason Zucker, Noémie Elhadad

Long-form clinical summarization of hospital admissions has real-world significance because of its potential to help both clinicians and patients. The faithfulness of summaries is critical to their safe usage in clinical…

Domain AdaptationFormSentence