Using statistical techniques and replication samples for imputation of metabolite missing values
Background: Data preparation, such as missing values imputation and transformation, is the first step in any data analysis and requires crucial attention. Particularly, analysis of metabolites demands more preparation since those small compounds have recently been measurable in large scales with mass spectrometry techniques. We introduce novel statistical techniques for metabolite missing values imputation by utilizing replication samples. Results: To understand the nature of the missing values using replication samples, we obtained the empirical distribution of missing values and observed that the rate of missing values is approximately distributed as uniform across the metabolite range. Therefore, the missing values cannot be imputed with the lowest values. Using the identified distribution, we illustrated a simulation study to find an optimal imputation approach for metabolites. Conclusions: We demonstrated that the missing values in metabolomic data sets might not be necessarily low value. After identification of the nature of missing values, we validated K nearest neighborhood as an optimal approach for imputation.
Code (0)
등록된 구현이 없습니다.
Tasks
ImputationMissing ValuesSimilar Papers 제목 키워드 기반
Multi-View Variational Autoencoder for Missing Value Imputation in Untargeted Metabolomics
Background: Missing data is a common challenge in mass spectrometry-based metabolomics, which can lead to biased and incomplete analyses. The integration of whole-genome sequencing (WGS) data with metabolomics data has e…
Data IntegrationImputationMissing ValuesMetabolomic profiles in Jamaican children with and without autism spectrum disorder
Autism spectrum disorder (ASD) is a complex neurodevelopmental condition with a wide range of behavioral and cognitive impairments. While genetic and environmental factors are known to contribute to its etiology, the und…
ImputationMissing ValuesGlycolytic pyruvate kinase moonlighting activities in DNA replication initiation and elongation
Cells have evolved a metabolic control of DNA replication to respond to a wide range of nutritional conditions. Accumulating data suggest that this poorly understood control depends, at least in part, on Central Carbon M…
MS2MetGAN: Latent-space adversarial training for metabolite-spectrum matching in MS/MS database search
Database search is a widely used approach for identifying metabolites from tandem mass spectra (MS/MS). In this strategy, an experimental spectrum is matched against a user-specified database of candidate metabolites, an…
When to Impute? Imputation before and during cross-validation
Cross-validation (CV) is a technique used to estimate generalization error for prediction models. For pipeline modeling algorithms (i.e. modeling procedures with multiple steps), it has been recommended the entire sequen…
ImputationMissing ValuesvalidVariable Selection