paper-with-me

Papers

Faithful Model Evaluation for Model-Based Metrics

2023-12-19 · Palash Goyal, Qian Hu, Rahul Gupta

Statistical significance testing is used in natural language processing (NLP) to determine whether the results of a study or experiment are likely to be due to chance or if they reflect a genuine relationship. A key step in significance testing is the estimation of confidence interval which is a function of sample variance. Sample variance calculation is straightforward when evaluating against ground truth. However, in many cases, a metric model is often used for evaluation. For example, to compare toxicity of two large language models, a toxicity classifier is used for evaluation. Existing works usually do not consider the variance change due to metric model errors, which can lead to wrong conclusions. In this work, we establish the mathematical foundation of significance testing for model-based metrics. With experiments on public benchmark datasets and a production system, we show that considering metric model errors to calculate sample variances for model-based metrics changes the conclusions in certain experiments.

📄 PDF Abstract BibTeX arXiv:2312.17254

Code (0)

등록된 구현이 없습니다.

Tasks

model

Similar Papers 제목 키워드 기반

Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth

2026-05-24 · Yoav Gur-Arieh, Ana Marasović, Mor Geva arxiv

Chains of thought (CoTs) have become central in interpreting and auditing behaviors of large language models. Yet growing evidence suggests that these traces often fail to faithfully represent the computations behind a m…

A Causal Lens for Evaluating Faithfulness Metrics

2025-02-26 · Kerem Zaman, Shashank Srivastava

Large Language Models (LLMs) offer natural language explanations as an alternative to feature attribution methods for model interpretability. However, despite their plausibility, they may not reflect the model's internal…

Decision MakingFact CheckingModel EditingObject Counting

ED-FAITH: Evaluating Dialogue Summarization on Faithfulness

2022-11-15 · Sicong Huang, Asli Celikyilmaz, Haoran Li

Abstractive summarization models typically generate content unfaithful to the input, thus highlighting the significance of evaluating the faithfulness of generated summaries. Most faithfulness metrics are only evaluated …

Abstractive Text SummarizationLanguage ModelingLanguage Modelling

Towards Spatially-Aware and Optimally Faithful Concept-Based Explanations

2025-04-15 · Shubham Kumar, Dwip Dalal, Narendra Ahuja

Post-hoc, unsupervised concept-based explanation methods (U-CBEMs) are a promising tool for generating semantic explanations of the decision-making processes in deep neural networks, having applications in both model imp…

Decision Making

A Comparative Study of Faithfulness Metrics for Model Interpretability Methods

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Interpretable methods to reveal the internal reasoning processes behind machine learning models have attracted increasing attention in recent years. To quantify the extent to which the identified interpretations truly re…

Decision Making