paper-with-me

홈 › Papers

Production Machine Learning Pipelines: Empirical Analysis and Optimization Opportunities

2021-03-30 · Doris Xin, Hui Miao, Aditya Parameswaran, Neoklis Polyzotis

Machine learning (ML) is now commonplace, powering data-driven applications in various organizations. Unlike the traditional perception of ML in research, ML production pipelines are complex, with many interlocking analytical components beyond training, whose sub-parts are often run multiple times on overlapping subsets of data. However, there is a lack of quantitative evidence regarding the lifespan, architecture, frequency, and complexity of these pipelines to understand how data management research can be used to make them more efficient, effective, robust, and reproducible. To that end, we analyze the provenance graphs of 3000 production ML pipelines at Google, comprising over 450,000 models trained, spanning a period of over four months, in an effort to understand the complexity and challenges underlying production ML. Our analysis reveals the characteristics, components, and topologies of typical industry-strength ML pipelines at various granularities. Along the way, we introduce a specialized data model for representing and reasoning about repeatedly run components in these ML pipelines, which we call model graphlets. We identify several rich opportunities for optimization, leveraging traditional data management ideas. We show how targeting even one of these opportunities, i.e., identifying and pruning wasted computation that does not translate to model deployment, can reduce wasted computation cost by 50% without compromising the model deployment cadence.

📄 PDF Abstract BibTeX arXiv:2103.16007

Code (0)

등록된 구현이 없습니다.

Tasks

BIG-bench Machine LearningManagement

Methods 이 논문이 사용한 방법론

Pruning 설명 없음

Similar Papers 제목 키워드 기반

Auto-Validate by-History: Auto-Program Data Quality Constraints to Validate Recurring Data Pipelines

2023-06-04 · Dezhan Tu, Yeye He, Weiwei Cui, Song Ge 외

Data pipelines are widely employed in modern enterprises to power a variety of Machine-Learning (ML) and Business-Intelligence (BI) applications. Crucially, these pipelines are \emph{recurring} (e.g., daily or hourly) in…

Exploring Data Pipelines through the Process Lens: a Reference Model forComputer Vision

2021-07-05 · Agathe Balayn, Bogdan Kulynych, Seda Guerses

Researchers have identified datasets used for training computer vision (CV) models as an important source of hazardous outcomes, and continue to examine popular CV datasets to expose their harms. These works tend to trea…

Layered TPOT: Speeding up Tree-based Pipeline Optimization

2018-01-18 · Pieter Gijsbers, Joaquin Vanschoren, Randal S. Olson

With the demand for machine learning increasing, so does the demand for tools which make it easier to use. Automated machine learning (AutoML) tools have been developed to address this need, such as the Tree-Based Pipeli…

Automated Feature EngineeringAutoMLBIG-bench Machine LearningHyperparameter Optimization

PRETZEL: Opening the Black Box of Machine Learning Prediction Serving Systems

2018-10-14 · Yunseong Lee, Alberto Scolari, Byung-Gon Chun, Marco Domenico Santambrogio 외

Machine Learning models are often composed of pipelines of transformations. While this design allows to efficiently execute single model components at training time, prediction serving has different requirements such as …

BIG-bench Machine LearningPrediction

Two-stage Optimization for Machine Learning Workflow

2019-07-01 · Alexandre Quemy

Machines learning techniques plays a preponderant role in dealing with massive amount of data and are employed in almost every possible domain. Building a high quality machine learning model to be deployed in production …

AutoMLBIG-bench Machine LearningMeta-LearningVocal Bursts Valence Prediction