paper-with-me

홈 › Papers

A Big Data Lake for Multilevel Streaming Analytics

2020-09-25 · Ruoran Liu, Haruna Isah, Farhana Zulkernine

Large organizations are seeking to create new architectures and scalable platforms to effectively handle data management challenges due to the explosive nature of data rarely seen in the past. These data management challenges are largely posed by the availability of streaming data at high velocity from various sources in multiple formats. The changes in data paradigm have led to the emergence of new data analytics and management architecture. This paper focuses on storing high volume, velocity and variety data in the raw formats in a data storage architecture called a data lake. First, we present our study on the limitations of traditional data warehouses in handling recent changes in data paradigms. We discuss and compare different open source and commercial platforms that can be used to develop a data lake. We then describe our end-to-end data lake design and implementation approach using the Hadoop Distributed File System (HDFS) on the Hadoop Data Platform (HDP). Finally, we present a real-world data lake development use case for data stream ingestion, staging, and multilevel streaming analytics which combines structured and unstructured data. This study can serve as a guide for individuals or organizations planning to implement a data lake solution for their use cases.

📄 PDF Abstract BibTeX arXiv:2009.12415

Code (0)

등록된 구현이 없습니다.

Tasks

Management

Similar Papers 제목 키워드 기반

A Scalable Framework for Multilevel Streaming Data Analytics using Deep Learning

2019-07-15 · Shihao Ge, Haruna Isah, Farhana Zulkernine, Shahzad Khan

The rapid growth of data in velocity, volume, value, variety, and veracity has enabled exciting new opportunities and presented big challenges for businesses of all types. Recently, there has been considerable interest i…

Sentiment Analysis

TAIJI: MCP-based Multi-Modal Data Analytics on Data Lakes

2025-05-16 · Chao Zhang, Shaolei Zhang, Quehuan Liu, Sibei Chen 외

The variety of data in data lakes presents significant challenges for data analytics, as data scientists must simultaneously analyze multi-modal data, including structured, semi-structured, and unstructured data. While L…

AI AgentMachine Unlearning

LakeMLB: Data Lake Machine Learning Benchmark

2026-02-11 · Feiyu Pan, Tianbin Zhang, Aoqian Zhang, Yu Sun 외 arxiv

Modern data lakes have emerged as foundational platforms for large-scale machine learning, enabling flexible storage of heterogeneous data and structured analytics through table-oriented abstractions. Despite their growi…

Data Augmentation

Real Time Analytics: Algorithms and Systems

2017-08-07 · Arun Kejariwal, Sanjeev Kulkarni, Karthik Ramasamy

Velocity is one of the 4 Vs commonly used to characterize Big Data. In this regard, Forrester remarked the following in Q3 2014: "The high velocity, white-water flow of data from innumerable real-time data sources such a…

Multilevel Analysis of Cryptocurrency News using RAG Approach with Fine-Tuned Mistral Large Language Model

2025-08-25 · Bohdan M. Pavlyshenko arxiv

In the paper, we consider multilevel multitask analysis of cryptocurrency news using a fine-tuned Mistral 7B large language model with retrieval-augmented generation (RAG). On the first level of analytics, the fine-tuned…