paper-with-me

Papers

Synth4bench: a framework for generating synthetic genomics data for the evaluation of tumor-only somatic variant calling algorithms

2024-03-08 · bioRxiv 2024 3 · Styliani-Christina Fragkouli, Nikos Pechlivanis, Anastasia Anastasiadou, Georgios Karakatsoulis, Aspasia Orfanou, Panagoula Kollia, Andreas Agathangelidis, Fotis Psomopoulos

Motivation Somatic variant calling algorithms are widely used to detect genomic alterations associated with cancer. Evaluating their performance, even though being crucial, can be challenging due to the lack of high-quality ground truth datasets. To address this issue, we developed a synthetic data generation framework for benchmarking these algorithms, focusing on the TP53 gene, utilizing the NEATv3.3 simulator. We thoroughly evaluated the performance of Mutect2, Freebayes, VarDict, VarScan2 and LoFreq and compared their results with our synthetic ground truth, while observing their behavior. Synth4bench attempts to shed light on the underlying principles of each variant caller by presenting them with data from a given range across the genomics data feature space and inspecting their response. Results Using synthetic dataset as ground truth provides an excellent approach for evaluating the performance of tumor-only somatic variant calling algorithms. Our findings are supported by an independent statistical analysis that was performed on the same data and output from all callers. Overall, synth4bench leverages the effort of benchmarking algorithms by offering the opportunity to utilize a generated ground truth dataset. This kind of framework is essential in the field of cancer genomics, where precision is an ultimate necessity, especially for variants of low frequency. In this context, our approach makes comparison of various algorithms transparent, straightforward and also enhances their comparability.

📄 PDF Abstract BibTeX

Code (1)

BiodataAnalysisGroup/synth4bench

Tasks

BenchmarkingSynthetic Data Generation

Similar Papers 제목 키워드 기반

Feedback GAN (FBGAN) for DNA: a Novel Feedback-Loop Architecture for Optimizing Protein Functions

2018-04-05 · Anvita Gupta, James Zou

Generative Adversarial Networks (GANs) represent an attractive and novel approach to generate realistic data, such as genes, proteins, or drugs, in synthetic biology. Here, we apply GANs to generate synthetic DNA sequenc…

Textomics: A Dataset for Genomics Data Summary Generation

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Summarizing biomedical discovery from genomics data using natural languages is an essential step in biomedical research but is mostly done manually. Here, we introduce Textomics, a novel dataset of genomics data descript…

Textomics: A Dataset for Genomics Data Summary Generation

2022-05-01 · ACL 2022 5 · Mu-Chun Wang, Zixuan Liu, Sheng Wang

Summarizing biomedical discovery from genomics data using natural languages is an essential step in biomedical research but is mostly done manually. Here, we introduce Textomics, a novel dataset of genomics data descript…

Inferring feature importance with uncertainties in high-dimensional data

2021-09-02 · Pål Vegard Johnsen, Inga Strümke, Signe Riemer-Sørensen, Andrew Thomas DeWan 외

Estimating feature importance is a significant aspect of explaining data-based models. Besides explaining the model itself, an equally relevant question is which features are important in the underlying data generating p…

Feature ImportanceVocal Bursts Intensity Prediction

Biologically-Informed Hybrid Membership Inference Attacks on Generative Genomic Models

2025-11-10 · Asia Belfiore, Jonathan Passerat-Palmbach, Dmitrii Usynin arxiv

The increased availability of genetic data has transformed genomics research, but raised many privacy concerns regarding its handling due to its sensitive nature. This work explores the use of language models (LMs) for t…