paper-with-me

홈 › Papers

XIMAGENET-12: An Explainable AI Benchmark Dataset for Model Robustness Evaluation

2023-10-12 · Qiang Li, Dan Zhang, Shengzhao Lei, Xun Zhao, Porawit Kamnoedboon, Weiwei Li, Junhao Dong, Shuyan Li

Despite the promising performance of existing visual models on public benchmarks, the critical assessment of their robustness for real-world applications remains an ongoing challenge. To bridge this gap, we propose an explainable visual dataset, XIMAGENET-12, to evaluate the robustness of visual models. XIMAGENET-12 consists of over 200K images with 15,410 manual semantic annotations. Specifically, we deliberately selected 12 categories from ImageNet, representing objects commonly encountered in practical life. To simulate real-world situations, we incorporated six diverse scenarios, such as overexposure, blurring, and color changes, etc. We further develop a quantitative criterion for robustness assessment, allowing for a nuanced understanding of how visual models perform under varying conditions, notably in relation to the background. We make the XIMAGENET-12 dataset and its corresponding code openly accessible at \url{https://sites.google.com/view/ximagenet-12/home}. We expect the introduction of the XIMAGENET-12 dataset will empower researchers to thoroughly evaluate the robustness of their visual models under challenging conditions.

📄 PDF Abstract BibTeX arXiv:2310.08182

Code (0)

등록된 구현이 없습니다.

Tasks

Classification

Similar Papers 제목 키워드 기반

Finding the right XAI method -- A Guide for the Evaluation and Ranking of Explainable AI Methods in Climate Science

2023-03-01 · Philine Bommer, Marlene Kretschmer, Anna Hedström, Dilyara Bareeva 외

Explainable artificial intelligence (XAI) methods shed light on the predictions of machine learning algorithms. Several different approaches exist and have already been applied in climate science. However, usually missin…

Explainable artificial intelligenceExplainable Artificial Intelligence (XAI)

ESGBench: A Benchmark for Explainable ESG Question Answering in Corporate Sustainability Reports

2025-11-20 · Sherine George, Nithish Saji arxiv

We present ESGBench, a benchmark dataset and evaluation framework designed to assess explainable ESG question answering systems using corporate sustainability reports. The benchmark consists of domain-grounded questions …

Question Answering

RAGBench: Explainable Benchmark for Retrieval-Augmented Generation Systems

2024-06-25 · Robert Friel, Masha Belyi, Atindriyo Sanyal

Retrieval-Augmented Generation (RAG) has become a standard architectural pattern for incorporating domain-specific knowledge into user-facing chat applications powered by Large Language Models (LLMs). RAG systems are cha…

BenchmarkingRAGRetrievalRetrieval-augmented Generation

ExCAM: Explainable Cultural Awareness Metrics

2026-05-28 · Christoph Leiter, Haiyue Song, Hour Kaing, Jin Tei 외 arxiv

Evaluating the cultural awareness of large language models is crucial to ensure the fairness of generated text and the generalizability of applications across the world. Recent benchmarks explore cultural goods like food…

Question AnsweringText Generation

Explainable Face Recognition

2020-08-03 · ECCV 2020 8 · Jonathan R. Williford, Brandon B. May, Jeffrey Byrne

Explainable face recognition is the problem of explaining why a facial matcher matches faces. In this paper, we provide the first comprehensive benchmark and baseline evaluation for explainable face recognition. We defin…

Face RecognitionTriplet