paper-with-me

Papers

Provenance Data in the Machine Learning Lifecycle in Computational Science and Engineering

2019-10-09 · Renan Souza, Leonardo Azevedo, Vítor Lourenço, Elton Soares, Raphael Thiago, Rafael Brandão, Daniel Civitarese, Emilio Vital Brazil, Marcio Moreno, Patrick Valduriez, Marta Mattoso, Renato Cerqueira, Marco A. S. Netto

Machine Learning (ML) has become essential in several industries. In Computational Science and Engineering (CSE), the complexity of the ML lifecycle comes from the large variety of data, scientists' expertise, tools, and workflows. If data are not tracked properly during the lifecycle, it becomes unfeasible to recreate a ML model from scratch or to explain to stakeholders how it was created. The main limitation of provenance tracking solutions is that they cannot cope with provenance capture and integration of domain and ML data processed in the multiple workflows in the lifecycle while keeping the provenance capture overhead low. To handle this problem, in this paper we contribute with a detailed characterization of provenance data in the ML lifecycle in CSE; a new provenance data representation, called PROV-ML, built on top of W3C PROV and ML Schema; and extensions to a system that tracks provenance from multiple workflows to address the characteristics of ML and CSE, and to allow for provenance queries with a standard vocabulary. We show a practical use in a real case in the Oil and Gas industry, along with its evaluation using 48 GPUs in parallel.

📄 PDF Abstract BibTeX arXiv:1910.04223

Code (0)

등록된 구현이 없습니다.

Tasks

BIG-bench Machine Learning

Similar Papers 제목 키워드 기반

Workflow Provenance in the Lifecycle of Scientific Machine Learning

2020-09-30 · Renan Souza, Leonardo G. Azevedo, Vítor Lourenço, Elton Soares 외

Machine Learning (ML) has already fundamentally changed several businesses. More recently, it has also been profoundly impacting the computational science and engineering domains, like geoscience, climate science, and he…

BIG-bench Machine Learning

From Data to Decision: Data-Centric Infrastructure for Reproducible ML in Collaborative eScience

2025-06-19 · Zhiwei Li, Carl Kesselman, Tran Huy Nguyen, Benjamin Yixing Xu 외

Reproducibility remains a central challenge in machine learning (ML), especially in collaborative eScience projects where teams iterate over data, features, and models. Current ML workflows are often dynamic yet fragment…

Operationalising Artificial Intelligence Bills of Materials (AIBOMs) for Verifiable AI Provenance and Lifecycle Assurance

2026-03-17 · Petar Radanliev, Omar Santos, Carsten Maple, Kay Atefi arxiv

Artificial Intelligence (AI) systems are increasingly dependent on complex, multi-layered software supply chains that introduce challenges for reproducibility, transparency, and security assurance. This study presents an…

Atlas: A Framework for ML Lifecycle Provenance & Transparency

2025-02-26 · Marcin Spoczynski, Marcela S. Melara, Sebastian Szyller

The rapid adoption of open source machine learning (ML) datasets and models exposes today's AI applications to critical risks like data poisoning and supply chain attacks across the ML lifecycle. With growing regulatory …

Data Poisoning

Spiral Model Technique For Data Science & Machine Learning Lifecycle

2025-10-08 · Rohith Mahadevan arxiv

Analytics play an important role in modern business. Companies adapt data science lifecycles to their culture to seek productivity and improve their competitiveness among others. Data science lifecycles are fairly an imp…