paper-with-me

Papers

Eval Factsheets: A Structured Framework for Documenting AI Evaluations

2025-12-03 · Florian Bordes, Candace Ross, Justine T Kao, Evangelia Spiliopoulou, Adina Williams arxiv

The rapid proliferation of benchmarks has created significant challenges in reproducibility, transparency, and informed decision-making. However, unlike datasets and models -- which benefit from structured documentation frameworks like Datasheets and Model Cards -- evaluation methodologies lack systematic documentation standards. We introduce Eval Factsheets, a structured, descriptive framework for documenting AI system evaluations through a comprehensive taxonomy and questionnaire-based approach. Our framework organizes evaluation characteristics across five fundamental dimensions: Context (Who made the evaluation and when?), Scope (What does it evaluate?), Structure (With what the evaluation is built?), Method (How does it work?) and Alignment (In what ways is it reliable/valid/robust?). We implement this taxonomy as a practical questionnaire spanning five sections with mandatory and recommended documentation elements. Through case studies on multiple benchmarks, we demonstrate that Eval Factsheets effectively captures diverse evaluation paradigms -- from traditional benchmarks to LLM-as-judge methodologies -- while maintaining consistency and comparability. We hope Eval Factsheets are incorporated into both existing and newly released evaluation frameworks and lead to more transparency and reproducibility.

📄 PDF Abstract BibTeX arXiv:2512.04062

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards Assuring EU AI Act Compliance and Adversarial Robustness of LLMs

2024-10-04 · Tomas Bueno Momcilovic, Beat Buesser, Giulio Zizzo, Mark Purcell 외

Large language models are prone to misuse and vulnerable to security threats, raising significant safety and security concerns. The European Union's Artificial Intelligence Act seeks to enforce AI robustness in certain c…

Adversarial Robustness

BenchmarkCards: Large Language Model and Risk Reporting

2024-10-16 · Anna Sokol, Nuno Moniz, Elizabeth Daly, Michael Hind 외

Large language models (LLMs) offer powerful capabilities but also introduce significant risks. One way to mitigate these risks is through comprehensive pre-deployment evaluations using benchmarks designed to test for spe…

FairnessLanguage ModelingLanguage ModellingLarge Language Model+1

A Methodology for Creating AI FactSheets

2020-06-24 · John Richards, David Piorkowski, Michael Hind, Stephanie Houde 외

As AI models and services are used in a growing number of highstakes areas, a consensus is forming around the need for a clearer record of how these models and services are developed to increase trust. Several proposals …

Improving LLM Performance Through Black-Box Online Tuning: A Case for Adding System Specs to Factsheets for Trusted AI

2026-03-11 · Yonas Atinafu, Henry Lin, Robin Cohen arxiv

In this paper, we present a novel black-box online controller that uses only end-to-end measurements over short segments, without internal instrumentation, and hill climbing to maximize goodput, defined as the throughput…

Towards an Accountable and Reproducible Federated Learning: A FactSheets Approach

2022-02-25 · Nathalie Baracaldo, Ali Anwar, Mark Purcell, Ambrish Rawat 외

Federated Learning (FL) is a novel paradigm for the shared training of models based on decentralized and private data. With respect to ethical guidelines, FL is promising regarding privacy, but needs to excel vis-\`a-vis…

EthicsFederated Learning