paper-with-me

홈 › Papers

Neural-based Modeling for Performance Tuning of Spark Data Analytics

2021-01-20 · Khaled Zaouk, Fei Song, Chenghao Lyu, Yanlei Diao

Cloud data analytics has become an integral part of enterprise business operations for data-driven insight discovery. Performance modeling of cloud data analytics is crucial for performance tuning and other critical operations in the cloud. Traditional modeling techniques fail to adapt to the high degree of diversity in workloads and system behaviors in this domain. In this paper, we bring recent Deep Learning techniques to bear on the process of automated performance modeling of cloud data analytics, with a focus on Spark data analytics as representative workloads. At the core of our work is the notion of learning workload embeddings (with a set of desired properties) to represent fundamental computational characteristics of different jobs, which enable performance prediction when used together with job configurations that control resource allocation and other system knobs. Our work provides an in-depth study of different modeling choices that suit our requirements. Results of extensive experiments reveal the strengths and limitations of different modeling methods, as well as superior performance of our best performing method over a state-of-the-art modeling tool for cloud analytics.

📄 PDF Abstract BibTeX arXiv:2101.08167

Code (1)

udao-modeling/code 공식 구현 tf

Tasks

Diversity

Similar Papers 제목 키워드 기반

Large-Scale News Classification using BERT Language Model: Spark NLP Approach

2021-07-14 · Kuncahyo Setyo Nugroho, Anantha Yullian Sukmadewa, Novanto Yudistira

The rise of big data analytics on top of NLP increases the computational burden for text processing at scale. The problems faced in NLP are very high dimensional text, so it takes a high computation resource. The MapRedu…

Language ModelingLanguage ModellingNews Classification

Mobile Big Data Analytics Using Deep Learning and Apache Spark

2016-02-23 · Mohammad Abu Alsheikh, Dusit Niyato, Shaowei Lin, Hwee-Pink Tan 외

The proliferation of mobile devices, such as smartphones and Internet of Things (IoT) gadgets, results in the recent mobile big data (MBD) era. Collecting MBD is unprofitable unless suitable analytics and learning method…

Activity RecognitionDeep Learning

Predictive Price-Performance Optimization for Serverless Query Processing

2021-12-16 · Rathijit Sen, Abhishek Roy, Alekh Jindal

We present an efficient, parametric modeling framework for predictive resource allocations, focusing on the amount of computational resources, that can optimize for a range of price-performance objectives for data analyt…

Benchmarking Apache Spark and Hadoop MapReduce on Big Data Classification

2022-09-21 · Taha Tekdogan, Ali Cakmak

Most of the popular Big Data analytics tools evolved to adapt their working environment to extract valuable information from a vast amount of unstructured data. The ability of data mining techniques to filter this helpfu…

BenchmarkingManagement

Big Data analytics. Three use cases with R, Python and Spark

2016-09-30 · Philippe Besse, Brendan Guillouet, Jean-Michel Loubes

Management and analysis of big data are systematically associated with a data distributed architecture in the Hadoop and now Spark frameworks. This article offers an introduction for statisticians to these technologies b…

Collaborative FilteringManagementregression