PRETZEL: Opening the Black Box of Machine Learning Prediction Serving Systems
Machine Learning models are often composed of pipelines of transformations. While this design allows to efficiently execute single model components at training time, prediction serving has different requirements such as low latency, high throughput and graceful performance degradation under heavy load. Current prediction serving systems consider models as black boxes, whereby prediction-time-specific optimizations are ignored in favor of ease of deployment. In this paper, we present PRETZEL, a prediction serving system introducing a novel white box architecture enabling both end-to-end and multi-model optimizations. Using production-like model pipelines, our experiments show that PRETZEL is able to introduce performance improvements over different dimensions; compared to state-of-the-art approaches PRETZEL is on average able to reduce 99th percentile latency by 5.5x while reducing memory footprint by 25x, and increasing throughput by 4.7x.
Code (0)
등록된 구현이 없습니다.
Tasks
BIG-bench Machine LearningPredictionSimilar Papers 제목 키워드 기반
Machine Learning Algorithms to Predict Chess960 Result and Develop Opening Themes
This work focuses on the analysis of Chess 960, also known as Fischer Random Chess, a variant of traditional chess where the starting positions of the pieces are randomized. The study aims to predict the game outcome usi…
PositionTraining Neural Nets to Achieve Audio-to-Score Translation: Opening the Black-Box
It is suggested that the task of audio-to-score translation offers an adequate testbed to investigate the division of labor between background knowledge and machine learning in the domain of audio pattern recognition, wi…
TranslationTowards Opening the Black Box of Neural Machine Translation: Source and Target Interpretations of the Transformer
In Neural Machine Translation (NMT), each token prediction is conditioned on the source sentence and the target prefix (what has been previously translated at a decoding step). However, previous work on interpretability …
DecoderMachine TranslationNMTSentence+1Thermo-responsive closing and reopening artificial Venus Flytrap utilizing shape memory elastomers
Despite their often perceived static and slow nature, some plants can move faster than the blink of an eye. The rapid snap closure motion of the Venus flytrap (Dionaea muscipula) has long captivated the interest of resea…
Coarse Semantic Injection for LLM-Conditioned Structured Indoor Prediction
Large language models (LLMs) have recently been used as structured decoders for indoor understanding from 3D point-token inputs. However, point cloud encoders often under-represent thin structural elements such as doors …