paper-with-me

Papers

Karasu: A Collaborative Approach to Efficient Cluster Configuration for Big Data Analytics

2023-08-22 · Dominik Scheinert, Philipp Wiesner, Thorsten Wittkopp, Lauritz Thamsen, Jonathan Will, Odej Kao

Selecting the right resources for big data analytics jobs is hard because of the wide variety of configuration options like machine type and cluster size. As poor choices can have a significant impact on resource efficiency, cost, and energy usage, automated approaches are gaining popularity. Most existing methods rely on profiling recurring workloads to find near-optimal solutions over time. Due to the cold-start problem, this often leads to lengthy and costly profiling phases. However, big data analytics jobs across users can share many common properties: they often operate on similar infrastructure, using similar algorithms implemented in similar frameworks. The potential in sharing aggregated profiling runs to collaboratively address the cold start problem is largely unexplored. We present Karasu, an approach to more efficient resource configuration profiling that promotes data sharing among users working with similar infrastructures, frameworks, algorithms, or datasets. Karasu trains lightweight performance models using aggregated runtime information of collaborators and combines them into an ensemble method to exploit inherent knowledge of the configuration search space. Moreover, Karasu allows the optimization of multiple objectives simultaneously. Our evaluation is based on performance data from diverse workload executions in a public cloud environment. We show that Karasu is able to significantly boost existing methods in terms of performance, search time, and cost, even when few comparable profiling runs are available that share only partial common characteristics with the target job.

📄 PDF Abstract BibTeX arXiv:2308.11792

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Comparing Native and Non-native English Speakers' Behaviors in Collaborative Writing through Visual Analytics

2025-02-25 · Yuexi Chen, Yimin Xiao, Kazi Tasnim Zinat, Naomi Yamashita 외

Understanding collaborative writing dynamics between native speakers (NS) and non-native speakers (NNS) is critical for enhancing collaboration quality and team inclusivity. In this paper, we partnered with communication…

Private Hierarchical Clustering and Efficient Approximation

2019-04-09 · Xianrui Meng, Dimitrios Papadopoulos, Alina Oprea, Nikos Triandopoulos

In collaborative learning, multiple parties contribute their datasets to jointly deduce global machine learning models for numerous predictive tasks. Despite its efficacy, this learning paradigm fails to encompass critic…

ClusteringPrivacy Preserving

Reproducible and Portable Big Data Analytics in the Cloud

2021-12-17 · Xin Wang, Pei Guo, Xingyan Li, Aryya Gangopadhyay 외

Cloud computing has become a major approach to help reproduce computational experiments. Yet there are still two main difficulties in reproducing batch based big data analytics (including descriptive and predictive analy…

Cloud ComputingCPUDescriptiveGPU

CanaryBench: Stress Testing Privacy Leakage in Cluster-Level Conversation Summaries

2026-01-25 · Deep Mehta arxiv

Aggregate analytics over conversational data are increasingly used for safety monitoring, governance, and product analysis in large language model systems. A common practice is to embed conversations, cluster them, and p…

Harnessing Transparent Learning Analytics for Individualized Support through Auto-detection of Engagement in Face-to-Face Collaborative Learning

2024-01-03 · Qi Zhou, Wannapon Suraworachet, Mutlu Cukurova

Using learning analytics to investigate and support collaborative learning has been explored for many years. Recently, automated approaches with various artificial intelligence approaches have provided promising results …