paper-with-me

홈 › Papers

Benchmarking XAI Explanations with Human-Aligned Evaluations

2024-11-04 · Rémi Kazmierczak, Steve Azzolin, Eloïse Berthier, Anna Hedström, Patricia Delhomme, Nicolas Bousquet, Goran Frehse, Massimiliano Mancini, Baptiste Caramiaux, Andrea Passerini, Gianni Franchi

In this paper, we introduce PASTA (Perceptual Assessment System for explanaTion of Artificial intelligence), a novel framework for a human-centric evaluation of XAI techniques in computer vision. Our first key contribution is a human evaluation of XAI explanations on four diverse datasets (COCO, Pascal Parts, Cats Dogs Cars, and MonumAI) which constitutes the first large-scale benchmark dataset for XAI, with annotations at both the image and concept levels. This dataset allows for robust evaluation and comparison across various XAI methods. Our second major contribution is a data-based metric for assessing the interpretability of explanations. It mimics human preferences, based on a database of human evaluations of explanations in the PASTA-dataset. With its dataset and metric, the PASTA framework provides consistent and reliable comparisons between XAI techniques, in a way that is scalable but still aligned with human evaluations. Additionally, our benchmark allows for comparisons between explanations across different modalities, an aspect previously unaddressed. Our findings indicate that humans tend to prefer saliency maps over other explanation types. Moreover, we provide evidence that human assessments show a low correlation with existing XAI metrics that are numerically simulated by probing the model.

📄 PDF Abstract BibTeX arXiv:2411.02470

Code (0)

등록된 구현이 없습니다.

Tasks

Benchmarking

Similar Papers 제목 키워드 기반

DreamBench++: A Human-Aligned Benchmark for Personalized Image Generation

2024-06-24 · Yuang Peng, Yuxin Cui, Haomiao Tang, Zekun Qi 외

Personalized image generation holds great promise in assisting humans in everyday work and life due to its impressive function in creatively generating personalized content. However, current evaluations either are automa…

BenchmarkingImage GenerationPersonalized Image Generation

GeoArena: Evaluating Open-World Geographic Reasoning in Large Vision-Language Models

2025-09-04 · Pengyue Jia, Yingyi Zhang, Xiangyu Zhao, Sharon Li arxiv

Geographic reasoning is a fundamental cognitive capability that requires models to infer plausible locations by synthesizing visual evidence with spatial world knowledge. Despite recent advances in large vision-language …

Metrics vs Surveys: An Analysis for Human-Aligned Benchmarking in Social Robot Navigation

2025-10-03 · Stefano Trepella, Mauro Martini, Noé Pérez-Higueras, Andrea Ostuni 외 arxiv

Social, also called human-aware, navigation is a key challenge for integrating mobile robots into human environments. The evaluation of such systems is complex, as factors such as comfort, safety, and legibility must be …

Robot Navigation

Mind the XAI Gap: A Human-Centered LLM Framework for Democratizing Explainable AI

2025-06-13 · Eva Paraschou, Ioannis Arapakis, Sofia Yfantidou, Sebastian Macaluso 외

Artificial Intelligence (AI) is rapidly embedded in critical decision-making systems, however their foundational ``black-box'' models require eXplainable AI (XAI) solutions to enhance transparency, which are mostly orien…

BenchmarkingIn-Context Learning

Grounded but Misleading: Evaluating Semantic Alignment in AI-Generated Security Explanations

2026-02-04 · Heajun An, Connor Ng, Sandesh Sharma Dulal, Junghwan Kim 외 arxiv

Online scams increasingly leverage fluent and context-aware social engineering strategies, creating growing demand for AI systems that explain why a message may be risky. However, explanations that cite detector-derived …