paper-with-me

Papers

Machine Learning Pipelines: Provenance, Reproducibility and FAIR Data Principles

2020-06-22 · Sheeba Samuel, Frank Löffler, Birgitta König-Ries

Machine learning (ML) is an increasingly important scientific tool supporting decision making and knowledge generation in numerous fields. With this, it also becomes more and more important that the results of ML experiments are reproducible. Unfortunately, that often is not the case. Rather, ML, similar to many other disciplines, faces a reproducibility crisis. In this paper, we describe our goals and initial steps in supporting the end-to-end reproducibility of ML pipelines. We investigate which factors beyond the availability of source code and datasets influence reproducibility of ML experiments. We propose ways to apply FAIR data practices to ML workflows. We present our preliminary results on the role of our tool, ProvBook, in capturing and comparing provenance of ML experiments and their reproducibility using Jupyter Notebooks.

📄 PDF Abstract BibTeX arXiv:2006.12117

Code (0)

등록된 구현이 없습니다.

Tasks

BIG-bench Machine LearningDecision Making

Similar Papers 제목 키워드 기반

Facilitating the sharing of electrophysiology data analysis results through in-depth provenance capture

2023-11-16 · Cristiano André Köhler, Danylo Ulianych, Sonja Grün, Stefan Decker 외

Scientific research demands reproducibility and transparency, particularly in data-intensive fields like electrophysiology. Electrophysiology data is typically analyzed using scripts that generate output files, including…

Transfer Learning

PIDSMaker: Building and Evaluating Provenance-based Intrusion Detection Systems

2026-01-30 · Tristan Bilot, Baoxiang Jiang, Thomas Pasquier arxiv

Recent provenance-based intrusion detection systems (PIDSs) have demonstrated strong potential for detecting advanced persistent threats (APTs) by applying machine learning to system provenance graphs. However, evaluatin…

Intrusion Detection

Astronomical Pipeline Provenance: A Use Case Evaluation

2021-09-22 · Michael A. C. Johnson, Marcus Paradies, Marta Dembska, Kristen Lackeos 외

In this decade astronomy is undergoing a paradigm shift to handle data from next generation observatories such as the Square Kilometre Array (SKA) or the Vera C. Rubin Observatory (LSST). Producing real time data streams…

Astronomy

Debugging Machine Learning Pipelines

2020-02-11 · Raoni Lourenço, Juliana Freire, Dennis Shasha

Machine learning tasks entail the use of complex computational pipelines to reach quantitative and qualitative conclusions. If some of the activities in a pipeline produce erroneous or uninformative outputs, the pipeline…

BIG-bench Machine Learning

Flurry: a Fast Framework for Reproducible Multi-layered Provenance Graph Representation Learning

2022-03-05 · Maya Kapoor, Joshua Melton, Michael Ridenhour, Mahalavanya Sriram 외

Complex heterogeneous dynamic networks like knowledge graphs are powerful constructs that can be used in modeling data provenance from computer systems. From a security perspective, these attributed graphs enable causali…

Anomaly DetectionGraph ClassificationGraph Representation LearningKnowledge Graphs+1