Designing Disaggregated Evaluations of AI Systems: Choices, Considerations, and Tradeoffs
Disaggregated evaluations of AI systems, in which system performance is assessed and reported separately for different groups of people, are conceptually simple. However, their design involves a variety of choices. Some of these choices influence the results that will be obtained, and thus the conclusions that can be drawn; others influence the impacts -- both beneficial and harmful -- that a disaggregated evaluation will have on people, including the people whose data is used to conduct the evaluation. We argue that a deeper understanding of these choices will enable researchers and practitioners to design careful and conclusive disaggregated evaluations. We also argue that better documentation of these choices, along with the underlying considerations and tradeoffs that have been made, will help others when interpreting an evaluation's results and conclusions.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Measuring Biological Capabilities and Risks of AI Agents
This paper addresses a rapidly emerging policy challenge: how to generate and interpret credible evidence about the biological capabilities and risks of AI scientists, or agentic AI systems capable of autonomously or col…
Assessing the Fairness of AI Systems: AI Practitioners' Processes, Challenges, and Needs for Support
Various tools and practices have been developed to support practitioners in identifying, assessing, and mitigating fairness-related harms caused by AI systems. However, prior research has highlighted gaps between the int…
FairnessFaster, Cheaper, Better: Multi-Objective Hyperparameter Optimization for LLM and RAG Systems
While Retrieval Augmented Generation (RAG) has emerged as a popular technique for improving Large Language Model (LLM) systems, it introduces a large number of choices, parameters and hyperparameters that must be made or…
Bayesian OptimizationHyperparameter OptimizationLanguage ModelingLanguage Modelling+3SureMap: Simultaneous Mean Estimation for Single-Task and Multi-Task Disaggregated Evaluation
Disaggregated evaluation -- estimation of performance of a machine learning model on different subpopulations -- is a core task when assessing performance and group-fairness of AI systems. A key challenge is that evaluat…
FairnessDesigning RF-Powered Battery-Less Electronic Shelf Labels With COTS Components
This paper presents a preliminary study exploring the feasibility of designing batteryless electronic shelf labels (ESLs) powered by radio frequency wireless power transfer using commercial off-the-shelf components. The …