paper-with-me

Papers

Quality Metrics for Transparent Machine Learning With and Without Humans In the Loop Are Not Correlated

2021-07-01 · Felix Biessmann, Dionysius Refiano

The field explainable artificial intelligence (XAI) has brought about an arsenal of methods to render Machine Learning (ML) predictions more interpretable. But how useful explanations provided by transparent ML methods are for humans remains difficult to assess. Here we investigate the quality of interpretable computer vision algorithms using techniques from psychophysics. In crowdsourced annotation tasks we study the impact of different interpretability approaches on annotation accuracy and task time. We compare these quality metrics with classical XAI, automated quality metrics. Our results demonstrate that psychophysical experiments allow for robust quality assessment of transparency in machine learning. Interestingly the quality metrics computed without humans in the loop did not provide a consistent ranking of interpretability methods nor were they representative for how useful an explanation was for humans. These findings highlight the potential of methods from classical psychophysics for modern machine learning applications. We hope that our results provide convincing arguments for evaluating interpretability in its natural habitat, human-ML interaction, if the goal is to obtain an authentic assessment of interpretability.

📄 PDF Abstract BibTeX arXiv:2107.02033

Code (0)

등록된 구현이 없습니다.

Tasks

BIG-bench Machine LearningExplainable artificial intelligenceExplainable Artificial Intelligence (XAI)

Similar Papers 제목 키워드 기반

A psychophysics approach for quantitative comparison of interpretable computer vision models

2019-11-24 · Felix Biessmann, Dionysius Irza Refiano

The field of transparent Machine Learning (ML) has contributed many novel methods aiming at better interpretability for computer vision and ML models in general. But how useful the explanations provided by transparent ML…

BIG-bench Machine Learning

Towards Explainable Evaluation Metrics for Machine Translation

2023-06-22 · Christoph Leiter, Piyawat Lertvittayakumjorn, Marina Fomicheva, Wei Zhao 외

Unlike classical lexical overlap metrics such as BLEU, most current evaluation metrics for machine translation (for example, COMET or BERTScore) are based on black-box large language models. They often achieve strong cor…

Machine TranslationTranslation

Towards Explainable Evaluation Metrics for Natural Language Generation

2022-03-21 · Christoph Leiter, Piyawat Lertvittayakumjorn, Marina Fomicheva, Wei Zhao 외

Unlike classical lexical overlap metrics such as BLEU, most current evaluation metrics (such as BERTScore or MoverScore) are based on black-box language models such as BERT or XLM-R. They often achieve strong correlation…

Machine TranslationText GenerationTranslationXLM-R

Transparent Human Evaluation for Image Captioning

2022-01-16 · ACL ARR January 2022 1 · Anonymous

We establish a rubric-based human evaluation protocol for image captioning models. Our scoring rubrics and their definitions are carefully developed based on machine- and human-generated captions on the MSCOCO dataset. E…

Image Captioning

Transparent Human Evaluation for Image Captioning

2021-11-17 · NAACL 2022 7 · Jungo Kasai, Keisuke Sakaguchi, Lavinia Dunagan, Jacob Morrison 외

We establish THumB, a rubric-based human evaluation protocol for image captioning models. Our scoring rubrics and their definitions are carefully developed based on machine- and human-generated captions on the MSCOCO dat…

Image Captioning