paper-with-me

홈 › Papers

The METRIC-framework for assessing data quality for trustworthy AI in medicine: a systematic review

2024-02-21 · Daniel Schwabe, Katinka Becker, Martin Seyferth, Andreas Klaß, Tobias Schäffter

The adoption of machine learning (ML) and, more specifically, deep learning (DL) applications into all major areas of our lives is underway. The development of trustworthy AI is especially important in medicine due to the large implications for patients' lives. While trustworthiness concerns various aspects including ethical, technical and privacy requirements, we focus on the importance of data quality (training/test) in DL. Since data quality dictates the behaviour of ML products, evaluating data quality will play a key part in the regulatory approval of medical AI products. We perform a systematic review following PRISMA guidelines using the databases PubMed and ACM Digital Library. We identify 2362 studies, out of which 62 records fulfil our eligibility criteria. From this literature, we synthesise the existing knowledge on data quality frameworks and combine it with the perspective of ML applications in medicine. As a result, we propose the METRIC-framework, a specialised data quality framework for medical training data comprising 15 awareness dimensions, along which developers of medical ML applications should investigate a dataset. This knowledge helps to reduce biases as a major source of unfairness, increase robustness, facilitate interpretability and thus lays the foundation for trustworthy AI in medicine. Incorporating such systematic assessment of medical datasets into regulatory approval processes has the potential to accelerate the approval of ML products and builds the basis for new standards.

📄 PDF Abstract BibTeX arXiv:2402.13635

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Focus 설명 없음
Library 설명 없음

Similar Papers 제목 키워드 기반

Revisiting Lexicon Evaluation in Unsupervised Word Discovery

2026-06-04 · Simon Malan, Danel Slabbert, Herman Kamper arxiv

Building a lexicon from discovered word-like units is a central goal in zero-resource speech processing. But do our evaluations provide a trustworthy indication of lexicon quality? A common metric, normalized edit distan…

U-Trustworthy Models.Reliability, Competence, and Confidence in Decision-Making

2024-01-04 · Ritwik Vashistha, Arya Farahi

With growing concerns regarding bias and discrimination in predictive models, the AI community has increasingly focused on assessing AI system trustworthiness. Conventionally, trustworthy AI literature relies on the prob…

Decision MakingModel SelectionPhilosophy

Scalable Utility-Aware Multiclass Calibration

2025-10-29 · Mahmoud Hegazy, Michael I. Jordan, Aymeric Dieuleveut arxiv

Ensuring that classifiers are well-calibrated, i.e., their predictions align with observed frequencies, is a minimal and fundamental requirement for classifiers to be viewed as trustworthy. Existing methods for assessing…

LaajMeter: A Framework for LaaJ Evaluation

2025-08-13 · Samuel Ackerman, Gal Amram, Ora Nova Fandina, Eitan Farchi 외 arxiv

Large Language Models (LLMs) are increasingly used as evaluators in natural language processing tasks, a paradigm known as LLM-as-a-Judge (LaaJ). The analysis of a LaaJ software, commonly refereed to as meta-evaluation, …

Code Translation

Assessing Trustworthiness of Autonomous Systems

2023-05-05 · Gregory Chance, Dhaminda B. Abeywickrama, Beckett LeClair, Owen Kerr 외

As Autonomous Systems (AS) become more ubiquitous in society, more responsible for our safety and our interaction with them more frequent, it is essential that they are trustworthy. Assessing the trustworthiness of AS is…