Variation in prediction accuracy due to randomness in data division and fair evaluation using interval estimation
This paper attempts to answer a "simple question" in building predictive models using machine learning algorithms. Although diagnostic and predictive models for various diseases have been proposed using data from large cohort studies and machine learning algorithms, challenges remain in their generalizability. Several causes for this challenge have been pointed out, and partitioning of the dataset with randomness is considered to be one of them. In this study, we constructed 33,600 diabetes diagnosis models with "initial state" dependent randomness using autoML (automatic machine learning framework) and open diabetes data, and evaluated their prediction accuracy. The results showed that the prediction accuracy had an initial state-dependent distribution. Since this distribution could follow a normal distribution, we estimated the expected interval of prediction accuracy using statistical interval estimation in order to fairly compare the accuracy of the prediction models.
Code (0)
등록된 구현이 없습니다.
Tasks
AutoMLDiagnosticPredictionSimilar Papers 제목 키워드 기반
Testing randomness for cancer risk
There are numerous stochastic models for cancer risk for a given tissue. Many rely on the following two hypotheses. 1. There is a fixed probability that a given cell division will eventually lead to a cancerous cell. 2. …
Beyond Point Estimate: Inferring Ensemble Prediction Variation from Neuron Activation Strength in Recommender Systems
Despite deep neural network (DNN)'s impressive prediction performance in various domains, it is well known now that a set of DNN models trained with the same model specification and the same data can produce very differe…
Model-based Reinforcement LearningPredictionRecommendation SystemsUniversal probability-free prediction
We construct universal prediction systems in the spirit of Popper's falsifiability and Kolmogorov complexity and randomness. These prediction systems do not depend on any statistical assumptions (but under the IID assump…
Conformal PredictionPredictionQuantifying Inherent Randomness in Machine Learning Algorithms
Most machine learning (ML) algorithms have several stochastic elements, and their performances are affected by these sources of randomness. This paper uses an empirical study to systematically examine the effects of two …
BIG-bench Machine LearningTowards a performance characteristic curve for model evaluation: an application in information diffusion prediction
The information diffusion prediction on social networks aims to predict future recipients of a message, with practical applications in marketing and social media. While different prediction models all claim to perform we…
MarketingPrediction