paper-with-me

홈 › Papers

Ensuring Dataset Quality for Machine Learning Certification

2020-11-03 · Sylvaine Picard, Camille Chapdelaine, Cyril Cappi, Laurent Gardes, Eric Jenn, Baptiste Lefèvre, Thomas Soumarmon

In this paper, we address the problem of dataset quality in the context of Machine Learning (ML)-based critical systems. We briefly analyse the applicability of some existing standards dealing with data and show that the specificities of the ML context are neither properly captured nor taken into ac-count. As a first answer to this concerning situation, we propose a dataset specification and verification process, and apply it on a signal recognition system from the railway domain. In addi-tion, we also give a list of recommendations for the collection and management of datasets. This work is one step towards the dataset engineering process that will be required for ML to be used on safety critical systems.

📄 PDF Abstract BibTeX arXiv:2011.01799

Code (0)

등록된 구현이 없습니다.

Tasks

BIG-bench Machine LearningManagement

Similar Papers 제목 키워드 기반

SIEVE: Towards Verifiable Certification for Code-datasets

2025-10-02 · Fatou Ndiaye Mbodji, El-hacen Diallo, Jordan Samhi, Kui Liu 외 arxiv

Code agents and empirical software engineering rely on public code datasets, yet these datasets lack verifiable quality guarantees. Static 'dataset cards' inform, but they are neither auditable nor do they offer statisti…

Scalable Certified Segmentation via Randomized Smoothing

2021-07-01 · Marc Fischer, Maximilian Baader, Martin Vechev

We present a new certification method for image and point cloud segmentation based on randomized smoothing. The method leverages a novel scalable algorithm for prediction and certification that correctly accounts for mul…

Point Cloud SegmentationSegmentation

RefCo and its Checker: Improving Language Documentation Corpora’s Reusability Through a Semi-Automatic Review Process

2022-06-01 · LREC 2022 6 · Herbert Lange, Jocelyn Aznar

The QUEST (QUality ESTablished) project aims at ensuring the reusability of audio-visual datasets (Wamprechtshammer et al., 2022) by devising quality criteria and curating processes. RefCo (Reference Corpora) is an initi…

Trusted Artificial Intelligence: Towards Certification of Machine Learning Applications

2021-03-31 · Philip Matthias Winter, Sebastian Eder, Johannes Weissenböck, Christoph Schwald 외

Artificial Intelligence is one of the fastest growing technologies of the 21st century and accompanies us in our daily lives when interacting with technical applications. However, reliance on such technical systems is cr…

BIG-bench Machine LearningEthics

Federated Learning with Blockchain-Enhanced Machine Unlearning: A Trustworthy Approach

2024-05-27 · Xuhan Zuo, Minghao Wang, Tianqing Zhu, Lefeng Zhang 외

With the growing need to comply with privacy regulations and respond to user data deletion requests, integrating machine unlearning into IoT-based federated learning has become imperative. Traditional unlearning methods,…

Federated LearningMachine UnlearningManagement