paper-with-me

Papers

EvalCards: A Framework for Standardized Evaluation Reporting

2025-11-05 · Ruchira Dhar, Danae Sanchez Villegas, Antonia Karamolegkou, Alice Schiavone, Yifei Yuan, Xinyi Chen, Jiaang Li, Stella Frank, Laura De Grazia, Monorama Swain, Stephanie Brandl, Daniel Hershcovich, Anders Søgaard, Desmond Elliott arxiv

Evaluation has long been a central concern in NLP, and transparent reporting practices are more critical than ever in today's landscape of rapidly released open-access models. Drawing on a survey of recent work on evaluation and documentation, we identify three persistent shortcomings in current reporting practices: reproducibility, accessibility, and governance. We argue that existing standardization efforts remain insufficient and introduce Evaluation Disclosure Cards (EvalCards) as a path forward. EvalCards are designed to enhance transparency for both researchers and practitioners while providing a practical foundation to meet emerging governance requirements.

📄 PDF Abstract BibTeX arXiv:2511.21695

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting

2026-06-08 · Avijit Ghosh, Anka Reuel, Jenny Chim, Wm. Matthew Kennedy 외 arxiv

AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blogs. The cost is interpretive: readers cannot reliably compare results across sour…

Scorecards for Synthetic Medical Data Evaluation and Reporting

2024-06-17 · Ghada Zamzmi, Adarsh Subbaswamy, Elena Sizikova, Edward Margerrison 외

Although interest in synthetic medical data (SMD) for training and testing AI methods is growing, the absence of a standardized framework to evaluate its quality and applicability hinders its wider adoption. Here, we out…

Position: Restructuring of Categories and Implementation of Guidelines Essential for VLM Adoption in Healthcare

2025-05-12 · Amara Tariq, Rimita Lahiri, Charles Kahn, Imon Banerjee

The intricate and multifaceted nature of vision language model (VLM) development, adaptation, and application necessitates the establishment of clear and standardized reporting protocols, particularly within the high-sta…

Position

A Standardized Framework For Evaluating Gene Expression Generative Models

2026-03-11 · Andrea Rubbi, Andrea Giuseppe Di Francesco, Mohammad Lotfollahi, Pietro Liò arxiv

The rapid development of generative models for single-cell gene expression data has created an urgent need for standardised evaluation frameworks. Current evaluation practices suffer from inconsistent metric implementati…

ELEVATE-GenAI: Reporting Guidelines for the Use of Large Language Models in Health Economics and Outcomes Research: an ISPOR Working Group on Generative AI Report

2024-12-23 · Rachael L. Fleurence, Dalia Dawoud, Jiang Bian, Mitchell K. Higashi 외

Introduction: Generative artificial intelligence (AI), particularly large language models (LLMs), holds significant promise for Health Economics and Outcomes Research (HEOR). However, standardized reporting guidance for …

FairnessSystematic Literature Review