paper-with-me

홈 › Papers

REIN: A Comprehensive Benchmark Framework for Data Cleaning Methods in ML Pipelines

2023-02-09 · Mohamed Abdelaal, Christian Hammacher, Harald Schoening

Nowadays, machine learning (ML) plays a vital role in many aspects of our daily life. In essence, building well-performing ML applications requires the provision of high-quality data throughout the entire life-cycle of such applications. Nevertheless, most of the real-world tabular data suffer from different types of discrepancies, such as missing values, outliers, duplicates, pattern violation, and inconsistencies. Such discrepancies typically emerge while collecting, transferring, storing, and/or integrating the data. To deal with these discrepancies, numerous data cleaning methods have been introduced. However, the majority of such methods broadly overlook the requirements imposed by downstream ML models. As a result, the potential of utilizing these data cleaning methods in ML pipelines is predominantly unrevealed. In this work, we introduce a comprehensive benchmark, called REIN1, to thoroughly investigate the impact of data cleaning methods on various ML models. Through the benchmark, we provide answers to important research questions, e.g., where and whether data cleaning is a necessary step in ML pipelines. To this end, the benchmark examines 38 simple and advanced error detection and repair methods. To evaluate these methods, we utilized a wide collection of ML models trained on 14 publicly-available datasets covering different domains and encompassing realistic as well as synthetic error profiles.

📄 PDF Abstract BibTeX arXiv:2302.04702

Code (1)

mohamedyd/rein-benchmark 공식 구현

Tasks

Missing Values

Methods 이 논문이 사용한 방법론

Repair 설명 없음

Similar Papers 제목 키워드 기반

Towards Practical Multi-Robot Hybrid Tasks Allocation for Autonomous Cleaning

2023-03-12 · Yabin Wang, Xiaopeng Hong, Zhiheng Ma, Tiedong Ma 외

Task allocation plays a vital role in multi-robot autonomous cleaning systems, where multiple robots work together to clean a large area. However, most current studies mainly focus on deterministic, single-task allocatio…

Deep Reinforcement Learning

Towards Robust Face Recognition with Comprehensive Search

2022-08-29 · Manyuan Zhang, Guanglu Song, Yu Liu, Hongsheng Li

Data cleaning, architecture, and loss function design are important factors contributing to high-performance face recognition. Previously, the research community tries to improve the performance of each single aspect but…

Face RecognitionRobust Face Recognition

Clean First, Align Later: Benchmarking Preference Data Cleaning for Reliable LLM Alignment

2025-09-28 · Samuel Yeh, Sharon Li arxiv

Human feedback plays a pivotal role in aligning large language models (LLMs) with human preferences. However, such feedback is often noisy or inconsistent, which can degrade the quality of reward models and hinder alignm…

Path Planning of Cleaning Robot with Reinforcement Learning

2022-08-17 · Woohyeon Moon, Bumgeun Park, Sarvar Hussain Nengroo, TaeYoung Kim 외

Recently, as the demand for cleaning robots has steadily increased, therefore household electricity consumption is also increasing. To solve this electricity consumption issue, the problem of efficient path planning for …

reinforcement-learningReinforcement LearningReinforcement Learning (RL)Transfer Learning

Data Cleaning and Machine Learning: A Systematic Literature Review

2023-10-03 · Pierre-Olivier Côté, Amin Nikanjam, Nafisa Ahmed, Dmytro Humeniuk 외

Context: Machine Learning (ML) is integrated into a growing number of systems for various applications. Because the performance of an ML model is highly dependent on the quality of the data it has been trained on, there …

ImputationOutlier DetectionSystematic Literature Review