paper-with-me

홈 › Papers

Subtle Data Crimes: Naively training machine learning algorithms could lead to overly-optimistic results

2021-09-16 · Efrat Shimron, Jonathan I. Tamir, Ke Wang, Michael Lustig

While open databases are an important resource in the Deep Learning (DL) era, they are sometimes used "off-label": data published for one task are used for training algorithms for a different one. This work aims to highlight that in some cases, this common practice may lead to biased, overly-optimistic results. We demonstrate this phenomenon for inverse problem solvers and show how their biased performance stems from hidden data preprocessing pipelines. We describe two preprocessing pipelines typical of open-access databases and study their effects on three well-established algorithms developed for Magnetic Resonance Imaging (MRI) reconstruction: Compressed Sensing (CS), Dictionary Learning (DictL), and DL. In this large-scale study we performed extensive computations. Our results demonstrate that the CS, DictL and DL algorithms yield systematically biased results when naively trained on seemingly-appropriate data: the Normalized Root Mean Square Error (NRMSE) improves consistently with the preprocessing extent, showing an artificial increase of 25%-48% in some cases. Since this phenomenon is generally unknown, biased results are sometimes published as state-of-the-art; we refer to that as subtle data crimes. This work hence raises a red flag regarding naive off-label usage of Big Data and reveals the vulnerability of modern inverse problem solvers to the resulting bias.

📄 PDF Abstract BibTeX arXiv:2109.08237

Code (2)

mikgroup/subtle_data_crimes 공식 구현 pytorch
mikgroup/subtle_inverse_crimes 공식 구현 pytorch

Tasks

compressed sensingDictionary LearningMRI Reconstruction

Similar Papers 제목 키워드 기반

Prediction of Homicides in Urban Centers: A Machine Learning Approach

2020-08-16 · José Ribeiro, Lair Meneses, Denis Costa, Wando Miranda 외

Relevant research has been highlighted in the computing community to develop machine learning models capable of predicting the occurrence of crimes, analyzing contexts of crimes, extracting profiles of individuals linked…

BIG-bench Machine Learning

Crime Topic Modeling

2017-01-05 · Da Kuang, P. Jeffrey Brantingham, Andrea L. Bertozzi

The classification of crime into discrete categories entails a massive loss of information. Crimes emerge out of a complex mix of behaviors and situations, yet most of these details cannot be captured by singular crime t…

Clustering

An Analysis Focused on Womens Safety: Can VAD Models Be Enhanced by a Multi-modal Dataset?

2026-05-25 · Sangeeta ., Maddikuntla Sai Prajwal, Debi Prosad Dogra, Kamalakar Vijay Thakare 외 arxiv

Women's safety and security are paramount for a modern society. Often, crimes scenes get recorded through low-resolution CCTV cameras limiting the efficiency of video anomaly detection (VAD) models. Despite substantial p…

Video Anomaly Detection

Quantum Algorithms: A New Frontier in Financial Crime Prevention

2024-03-27 · Abraham Itzhak Weinberg, Alessio Faccia

Financial crimes fast proliferation and sophistication require novel approaches that provide robust and effective solutions. This paper explores the potential of quantum algorithms in combating financial crimes. It highl…

ManagementQuantum Machine Learning

Comprehensive Botnet Detection by Mitigating Adversarial Attacks, Navigating the Subtleties of Perturbation Distances and Fortifying Predictions with Conformal Layers

2024-09-01 · Rahul Yumlembam, Biju Issac, Seibu Mary Jacob, Longzhi Yang

Botnets are computer networks controlled by malicious actors that present significant cybersecurity challenges. They autonomously infect, propagate, and coordinate to conduct cybercrimes, necessitating robust detection m…

Conformal PredictionGenerative Adversarial Network