Synth4bench: a framework for generating synthetic genomics data for the evaluation of tumor-only somatic variant calling algorithms
Motivation Somatic variant calling algorithms are widely used to detect genomic alterations associated with cancer. Evaluating their performance, even though being crucial, can be challenging due to the lack of high-quality ground truth datasets. To address this issue, we developed a synthetic data generation framework for benchmarking these algorithms, focusing on the TP53 gene, utilizing the NEATv3.3 simulator. We thoroughly evaluated the performance of Mutect2, Freebayes, VarDict, VarScan2 and LoFreq and compared their results with our synthetic ground truth, while observing their behavior. Synth4bench attempts to shed light on the underlying principles of each variant caller by presenting them with data from a given range across the genomics data feature space and inspecting their response. Results Using synthetic dataset as ground truth provides an excellent approach for evaluating the performance of tumor-only somatic variant calling algorithms. Our findings are supported by an independent statistical analysis that was performed on the same data and output from all callers. Overall, synth4bench leverages the effort of benchmarking algorithms by offering the opportunity to utilize a generated ground truth dataset. This kind of framework is essential in the field of cancer genomics, where precision is an ultimate necessity, especially for variants of low frequency. In this context, our approach makes comparison of various algorithms transparent, straightforward and also enhances their comparability.
Code (1)
Tasks
BenchmarkingSynthetic Data GenerationSimilar Papers 제목 키워드 기반
Feedback GAN (FBGAN) for DNA: a Novel Feedback-Loop Architecture for Optimizing Protein Functions
Generative Adversarial Networks (GANs) represent an attractive and novel approach to generate realistic data, such as genes, proteins, or drugs, in synthetic biology. Here, we apply GANs to generate synthetic DNA sequenc…
Textomics: A Dataset for Genomics Data Summary Generation
Summarizing biomedical discovery from genomics data using natural languages is an essential step in biomedical research but is mostly done manually. Here, we introduce Textomics, a novel dataset of genomics data descript…
Textomics: A Dataset for Genomics Data Summary Generation
Summarizing biomedical discovery from genomics data using natural languages is an essential step in biomedical research but is mostly done manually. Here, we introduce Textomics, a novel dataset of genomics data descript…
Inferring feature importance with uncertainties in high-dimensional data
Estimating feature importance is a significant aspect of explaining data-based models. Besides explaining the model itself, an equally relevant question is which features are important in the underlying data generating p…
Feature ImportanceVocal Bursts Intensity PredictionBiologically-Informed Hybrid Membership Inference Attacks on Generative Genomic Models
The increased availability of genetic data has transformed genomics research, but raised many privacy concerns regarding its handling due to its sensitive nature. This work explores the use of language models (LMs) for t…