paper-with-me

홈 › Papers

Auto-Tag: Tagging-Data-By-Example in Data Lakes

2021-12-11 · Yeye He, Jie Song, Yue Wang, Surajit Chaudhuri, Vishal Anil, Blake Lassiter, Yaron Goland, Gaurav Malhotra

As data lakes become increasingly popular in large enterprises today, there is a growing need to tag or classify data assets (e.g., files and databases) in data lakes with additional metadata (e.g., semantic column-types), as the inferred metadata can enable a range of downstream applications like data governance (e.g., GDPR compliance), and dataset search. Given the sheer size of today's enterprise data lakes with petabytes of data and millions of data assets, it is imperative that data assets can be `auto-tagged'', using lightweight inference algorithms and minimal user input. In this work, we develop Auto-Tag, a corpus-driven approach that automates data-tagging of \textit{custom} data types in enterprise data lakes. Using Auto-Tag, users only need to provide \textit{one} example column to demonstrate the desired data-type to tag. Leveraging an index structure built offline using a lightweight scan of the data lake, which is analogous to pre-training in machine learning, Auto-Tag can infer suitable data patterns to best describe'' the underlying domain'' of the given column at an interactive speed, which can then be used to tag additional data of the same type'' in data lakes. The Auto-Tag approach can adapt to custom data-types, and is shown to be both accurate and efficient. Part of Auto-Tag ships as a `custom-classification'' feature in a cloud-based data governance and catalog solution \textit{Azure Purview}.

📄 PDF Abstract BibTeX arXiv:2112.06049

Code (0)

등록된 구현이 없습니다.

Tasks

TAG

Similar Papers 제목 키워드 기반

Retrieve, Merge, Predict: Augmenting Tables with Data Lakes

2024-02-09 · Riccardo Cappuzzo, Aimee Coelho, Felix Lefebvre, Paolo Papotti 외

Machine-learning from a disparate set of tables, a data lake, requires assembling features by merging and aggregating tables. Data discovery can extend autoML to data tables by automating these steps. We present an in-de…

AutoMLBenchmarkingFeature Engineering

Identification and classification of exfoliated graphene flakes from microscopy images using a hierarchical deep convolutional neural network

2022-03-29 · Soroush Mahjoubi, Fan Ye, Yi Bao, Weina Meng 외

Identification of the mechanically exfoliated graphene flakes and classification of the thickness is important in the nanomanufacturing of next-generation materials and devices that overcome the bottleneck of Moore's Law…

Deep Learning

Machine Learning-based Automatic Graphene Detection with Color Correction for Optical Microscope Images

2021-03-24 · Hui-Ying Siao, Siyu Qi, Zhi Ding, Chia-Yu Lin 외

Graphene serves critical application and research purposes in various fields. However, fabricating high-quality and large quantities of graphene is time-consuming and it requires heavy human resource labor costs. In this…

BIG-bench Machine LearningImage SegmentationSemantic Segmentation

Constructing a High Temporal Resolution Global Lakes Dataset via Swin-Unet with Applications to Area Prediction

2024-08-20 · Yutian Han, Baoxiang Huang, He Gao

Lakes provide a wide range of valuable ecosystem services, such as water supply, biodiversity habitats, and carbon sequestration. However, lakes are increasingly threatened by climate change and human activities. Therefo…

Attribute

Targeted Semantic Segmentation of Himalayan Glacial Lakes Using Time-Series SAR: Towards Automated GLOF Early Warning

2025-12-30 · Pawan Adhikari, Satish Raj Regmi, Hari Ram Shrestha arxiv

Glacial Lake Outburst Floods (GLOFs) are one of the most devastating climate change induced hazards. Existing remote monitoring approaches often prioritise maximising spatial coverage to train generalistic models or rely…

Semantic Segmentation