paper-with-me

Papers

OpenHEXAI: An Open-Source Framework for Human-Centered Evaluation of Explainable Machine Learning

2024-02-20 · Jiaqi Ma, Vivian Lai, Yiming Zhang, Chacha Chen, Paul Hamilton, Davor Ljubenkov, Himabindu Lakkaraju, Chenhao Tan

Recently, there has been a surge of explainable AI (XAI) methods driven by the need for understanding machine learning model behaviors in high-stakes scenarios. However, properly evaluating the effectiveness of the XAI methods inevitably requires the involvement of human subjects, and conducting human-centered benchmarks is challenging in a number of ways: designing and implementing user studies is complex; numerous design choices in the design space of user study lead to problems of reproducibility; and running user studies can be challenging and even daunting for machine learning researchers. To address these challenges, this paper presents OpenHEXAI, an open-source framework for human-centered evaluation of XAI methods. OpenHEXAI features (1) a collection of diverse benchmark datasets, pre-trained models, and post hoc explanation methods; (2) an easy-to-use web application for user study; (3) comprehensive evaluation metrics for the effectiveness of post hoc explanation methods in the context of human-AI decision making tasks; (4) best practice recommendations of experiment documentation; and (5) convenient tools for power analysis and cost estimation. OpenHEAXI is the first large-scale infrastructural effort to facilitate human-centered benchmarks of XAI methods. It simplifies the design and implementation of user studies for XAI methods, thus allowing researchers and practitioners to focus on the scientific questions. Additionally, it enhances reproducibility through standardized designs. Based on OpenHEXAI, we further conduct a systematic benchmark of four state-of-the-art post hoc explanation methods and compare their impacts on human-AI decision making tasks in terms of accuracy, fairness, as well as users' trust and understanding of the machine learning model.

📄 PDF Abstract BibTeX arXiv:2403.05565

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingFairness

Methods 이 논문이 사용한 방법론

Focus 설명 없음
HOC 설명 없음

Similar Papers 제목 키워드 기반

HumaniBench: A Human-Centric Framework for Large Multimodal Models Evaluation

2025-05-16 · Shaina Raza, Aravind Narayanan, Vahid Reza Khazaie, Ashmal Vayani 외

Large multimodal models (LMMs) now excel on many vision language benchmarks, however, they still struggle with human centered criteria such as fairness, ethics, empathy, and inclusivity, key to aligning with human values…

BenchmarkingEthicsFairnessQuestion Answering+3

HELM: A Human-Centered Evaluation Framework for LLM-Powered Recommender Systems

2026-01-27 · Sushant Mehta arxiv

The integration of Large Language Models (LLMs) into recommendation systems has introduced unprecedented capabilities for natural language understanding, explanation generation, and conversational interactions. However, …

Natural Language UnderstandingCollaborative FilteringExplanation GenerationRecommendation Systems

Patient-Centered Summarization Framework for AI Clinical Summarization: A Mixed-Methods Design

2025-10-31 · Maria Lizarazo Jimenez, Ana Gabriela Claros, Kieran Green, David Toro-Tobon 외 arxiv

Large Language Models (LLMs) are increasingly demonstrating the potential to reach human-level performance in generating clinical summaries from patient-clinician conversations. However, these summaries often focus on pa…

When Words Outperform Vision: VLMs Can Self-Improve Via Text-Only Training For Human-Centered Decision Making

2025-03-21 · Zhe Hu, Jing Li, Yu Yin

Embodied decision-making is fundamental for AI agents operating in real-world environments. While Visual Language Models (VLMs) have advanced this capability, they still struggle with complex decisions, particularly in h…

Decision Making

OmniForce: On Human-Centered, Large Model Empowered and Cloud-Edge Collaborative AutoML System

2023-03-01 · Chao Xue, Wei Liu, Shuai Xie, Zhenfang Wang 외

Automated machine learning (AutoML) seeks to build ML models with minimal human effort. While considerable research has been conducted in the area of AutoML in general, aiming to take humans out of the loop when building…

AutoML