paper-with-me

홈 › Papers

AutoCure: Automated Tabular Data Curation Technique for ML Pipelines

2023-04-26 · Mohamed Abdelaal, Rashmi Koparde, Harald Schoening

Machine learning algorithms have become increasingly prevalent in multiple domains, such as autonomous driving, healthcare, and finance. In such domains, data preparation remains a significant challenge in developing accurate models, requiring significant expertise and time investment to search the huge search space of well-suited data curation and transformation tools. To address this challenge, we present AutoCure, a novel and configuration-free data curation pipeline that improves the quality of tabular data. Unlike traditional data curation methods, AutoCure synthetically enhances the density of the clean data fraction through an adaptive ensemble-based error detection method and a data augmentation module. In practice, AutoCure can be integrated with open source tools, e.g., Auto-sklearn, H2O, and TPOT, to promote the democratization of machine learning. As a proof of concept, we provide a comparative evaluation of AutoCure against 28 combinations of traditional data curation tools, demonstrating superior performance and predictive accuracy without user intervention. Our evaluation shows that AutoCure is an effective approach to automating data preparation and improving the accuracy of machine learning models.

📄 PDF Abstract BibTeX arXiv:2304.13636

Code (1)

mohamedyd/AutoCure 공식 구현

Tasks

Autonomous DrivingData Augmentation

Similar Papers 제목 키워드 기반

Augmented Understanding and Automated Adaptation of Curation Rules

2020-07-17 · Alireza Tabebordbar

Over the past years, there has been many efforts to curate and increase the added value of the raw data. Data curation has been defined as activities and processes an analyst undertakes to transform the raw data into con…

Entity Extraction using GANPOS

Text Serialization and Their Relationship with the Conventional Paradigms of Tabular Machine Learning

2024-06-19 · Kyoka Ono, Simon A. Lee

Recent research has explored how Language Models (LMs) can be used for feature representation and prediction in tabular machine learning tasks. This involves employing text serialization and supervised fine-tuning (SFT) …

TabDPT: Scaling Tabular Foundation Models

2024-10-23 · Junwei Ma, Valentin Thomas, Rasa Hosseinzadeh, Hamidreza Kamkari 외

The challenges faced by neural networks on tabular data are well-documented and have hampered the progress of tabular foundation models. Techniques leveraging in-context learning (ICL) have shown promise here, allowing f…

In-Context LearningSelf-Supervised Learning

Initial Exploration of Zero-Shot Privacy Utility Tradeoffs in Tabular Data Using GPT-4

2024-04-07 · Bishwas Mandal, George Amariucai, Shuangqing Wei

We investigate the application of large language models (LLMs), specifically GPT-4, to scenarios involving the tradeoff between privacy and utility in tabular data. Our approach entails prompting GPT-4 by transforming ta…

Fairness

Curated LLM: Synergy of LLMs and Data Curation for tabular augmentation in low-data regimes

2023-12-19 · Nabeel Seedat, Nicolas Huynh, Boris van Breugel, Mihaela van der Schaar

Machine Learning (ML) in low-data settings remains an underappreciated yet crucial problem. Hence, data augmentation methods to increase the sample size of datasets needed for ML are key to unlocking the transformative p…

Data Augmentation