paper-with-me

Papers

A Fast, Scalable, Universal Approach For Distributed Data Aggregations

2020-10-27 · Niranda Perera, Vibhatha Abeykoon, Chathura Widanage, Supun Kamburugamuve, Thejaka Amila Kanewala, Pulasthi Wickramasinghe, Ahmet Uyar, Hasara Maithree, Damitha Lenadora, Geoffrey Fox

In the current era of Big Data, data engineering has transformed into an essential field of study across many branches of science. Advancements in Artificial Intelligence (AI) have broadened the scope of data engineering and opened up new applications in both enterprise and research communities. Aggregations (also termed reduce in functional programming) are an integral functionality in these applications. They are traditionally aimed at generating meaningful information on large data-sets, and today, they are being used for engineering more effective features for complex AI models. Aggregations are usually carried out on top of data abstractions such as tables/ arrays and are combined with other operations such as grouping of values. There are frameworks that excel in the said domains individually. But, we believe that there is an essential requirement for a data analytics tool that can universally integrate with existing frameworks, and thereby increase the productivity and efficiency of the entire data analytics pipeline. Cylon endeavors to fulfill this void. In this paper, we present Cylon's fast and scalable aggregation operations implemented on top of a distributed in-memory table structure that universally integrates with existing frameworks.

📄 PDF Abstract BibTeX arXiv:2010.14596

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Understanding and Comparing Scalable Gaussian Process Regression for Big Data

2018-11-03 · Haitao Liu, Jianfei Cai, Yew-Soon Ong, Yi Wang

As a non-parametric Bayesian model which produces informative predictive distribution, Gaussian process (GP) has been widely used in various fields, like regression, classification and optimization. The cubic complexity …

regression

Generalized Robust Bayesian Committee Machine for Large-scale Gaussian Process Regression

2018-06-03 · ICML 2018 7 · Haitao Liu, Jianfei Cai, Yi Wang, Yew-Soon Ong

In order to scale standard Gaussian process (GP) regression to large-scale datasets, aggregation models employ factorized training process and then combine predictions from distributed experts. The state-of-the-art aggre…

Distributed Computingregression

Decentralized Event-Triggered Federated Learning with Heterogeneous Communication Thresholds

2022-04-07 · Shahryar Zehtabi, Seyyedali Hosseinalipour, Christopher G. Brinton

A recent emphasis of distributed learning research has been on federated learning (FL), in which model training is conducted by the data-collecting devices. Existing research on FL has mostly focused on a star topology l…

Federated Learning

Spherical Cap Packing Asymptotics and Rank-Extreme Detection

2015-11-19 · Kai Zhang

We study the spherical cap packing problem with a probabilistic approach. Such probabilistic considerations result in an asymptotic sharp universal uniform bound on the maximal inner product between any set of unit vecto…

A note about the generalisation of the C-tests

2014-12-30 · Jose Hernandez-Orallo

In this exploratory note we ask the question of what a measure of performance for all tasks is like if we use a weighting of tasks based on a difficulty function. This difficulty function depends on the complexity of the…