paper-with-me

홈 › Papers

Measuring Swampiness: Quantifying Chaos in Large Heterogeneous Data Repositories

2018-10-13 · Jung Luann, Whitaker Brendan, Chard Kyle, Elmore Aaron

As scientific data repositories and filesystems grow in size and complexity, they become increasingly disorganized. The coupling of massive quantities of data with poor organization makes it challenging for scientists to locate and utilize relevant data, thus slowing the process of analyzing data of interest. To address these issues, we explore an automated clustering approach for quantifying the organization of data repositories. Our parallel pipeline processes heterogeneous filetypes (e.g., text and tabular data), automatically clusters files based on content and metadata similarities, and computes a novel "cleanliness" score from the resulting clustering. We demonstrate the generation and accuracy of our cleanliness measure using both synthetic and real datasets, and conclude that it is more consistent than other potential cleanliness measures.

📄 PDF Abstract BibTeX arXiv:1810.05784

Code (0)

등록된 구현이 없습니다.

Tasks

Clustering

Similar Papers 제목 키워드 기반

Quantifying the Chaos Level of Infants' Environment via Unsupervised Learning

2019-12-10 · Priyanka Khante, Mai Lee Chang, Domingo Martinez, Kaya de Barbaro 외

Acoustic environments vary dramatically within the home setting. They can be a source of comfort and tranquility or chaos that can lead to less optimal cognitive development in children. Research to date has only subject…

BIG-bench Machine LearningClustering

Random Heterogeneous Neurochaos Learning Architecture for Data Classification

2024-10-30 · Remya Ajai A S, Nithin Nagaraj

Inspired by the human brain's structure and function, Artificial Neural Networks (ANN) were developed for data classification. However, existing Neural Networks, including Deep Neural Networks, do not mimic the brain's r…

Classification

Chaos and noise in evolutionary game dynamics

2025-03-28 · Maria Alejandra Ramirez, George Datseris, Arne Traulsen

Evolutionary game theory has traditionally employed deterministic models to describe population dynamics. These models, due to their inherent nonlinearities, can exhibit deterministic chaos, where population fluctuations…

The Multiscale Single-Index Model: A Stylized Model for Hierarchical Feature Learning

2026-07-03 · Joan Bruna arxiv

We consider the Multiscale Single-Index Model (MSIM), first introduced in \cite{oymak2021learning}, as a stylized model for hierarchical learning with \emph{scale separation}. Each layer extracts a shared single-index fe…

Measuring Heterogeneity in Machine Learning with Distributed Energy Distance

2025-01-27 · Mengchen Fan, Baocheng Geng, Roman Shterenberg, Joseph A. Casey 외

In distributed and federated learning, heterogeneity across data sources remains a major obstacle to effective model aggregation and convergence. We focus on feature heterogeneity and introduce energy distance as a sensi…

Federated Learning