paper-with-me

Papers

"Model Cards for Model Reporting" in 2024: Reclassifying Category of Ethical Considerations in Terms of Trustworthiness and Risk Management

2024-02-15 · DeBrae Kennedy-Mayo, Jake Gord

In 2019, the paper entitled "Model Cards for Model Reporting" introduced a new tool for documenting model performance and encouraged the practice of transparent reporting for a defined list of categories. One of the categories detailed in that paper is ethical considerations, which includes the subcategories of data, human life, mitigations, risks and harms, and use cases. We propose to reclassify this category in the original model card due to the recent maturing of the field known as trustworthy AI, a term which analyzes whether the algorithmic properties of the model indicate that the AI system is deserving of trust from its stakeholders. In our examination of trustworthy AI, we highlight three respected organizations - the European Commission's High-Level Expert Group on AI, the OECD, and the U.S.-based NIST - that have written guidelines on various aspects of trustworthy AI. These recent publications converge on numerous characteristics of the term, including accountability, explainability, fairness, privacy, reliability, robustness, safety, security, and transparency, while recognizing that the implementation of trustworthy AI varies by context. Our reclassification of the original model-card category known as ethical considerations involves a two-step process: 1) adding a new category known as trustworthiness, where the subcategories will be derived from the discussion of trustworthy AI in our paper, and 2) maintaining the subcategories of ethical considerations under a renamed category known as risk environment and risk management, a title which we believe better captures today's understanding of the essence of these topics. We hope that this reclassification will further the goals of the original paper and continue to prompt those releasing trained models to accompany these models with documentation that will assist in the evaluation of their algorithmic properties.

📄 PDF Abstract BibTeX arXiv:2403.15394

Code (0)

등록된 구현이 없습니다.

Tasks

FairnessManagementmodel

Similar Papers 제목 키워드 기반

Gender as a Variable in Natural-Language Processing: Ethical Considerations

2017-04-01 · WS 2017 4 · Brian Larson

Researchers and practitioners in natural-language processing (NLP) and related fields should attend to ethical principles in study design, ascription of categories/variables to study participants, and reporting of findin…

EvalCards: A Framework for Standardized Evaluation Reporting

2025-11-05 · Ruchira Dhar, Danae Sanchez Villegas, Antonia Karamolegkou, Alice Schiavone 외 arxiv

Evaluation has long been a central concern in NLP, and transparent reporting practices are more critical than ever in today's landscape of rapidly released open-access models. Drawing on a survey of recent work on evalua…

DREAMS: A python framework to train deep learning models with model card reporting for medical and health applications

2024-09-26 · Rabindra Khadka, Pedro G Lind, Anis Yazidi, Asma Belhadi

Electroencephalography (EEG) data provides a non-invasive method for researchers and clinicians to observe brain activity in real time. The integration of deep learning techniques with EEG data has significantly improved…

Deep LearningEEG

Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI

2022-04-03 · Mahima Pushkarna, Andrew Zaldivar, Oddur Kjartansson

As research and industry moves towards large-scale models capable of numerous downstream tasks, the complexity of understanding multi-modal datasets that give nuance to models rapidly increases. A clear and thorough unde…

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting

2026-06-08 · Avijit Ghosh, Anka Reuel, Jenny Chim, Wm. Matthew Kennedy 외 arxiv

AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blogs. The cost is interpretive: readers cannot reliably compare results across sour…