paper-with-me

홈 › Papers

Estimation of genome size using k-mer frequencies from corrected long reads

2020-03-26 · Hengchao Wang, Bo Liu, Yan Zhang, Fan Jiang, Yuwei Ren, Lijuan Yin, Hangwei Liu, Sen Wang, Wei Fan

The third-generation long reads sequencing technologies, such as PacBio and Nanopore, have great advantages over second-generation Illumina sequencing in de novo assembly studies. However, due to the inherent low base accuracy, third-generation sequencing data cannot be used for k-mer counting and estimating genomic profile based on k-mer frequencies. Thus, in current genome projects, second-generation data is also necessary for accurately determining genome size and other genomic characteristics. We show that corrected third-generation data can be used to count k-mer frequencies and estimate genome size reliably, in replacement of using second-generation data. Therefore, future genome projects can depend on only one sequencing technology to finish both assembly and k-mer analysis, which will largely decrease sequencing cost in both time and money. Moreover, we present a fast light-weight tool kmerfreq and use it to perform all the k-mer counting tasks in this work. We have demonstrated that corrected third-generation sequencing data can be used to estimate genome size and developed a new open-source C/C++ k-mer counting tool, kmerfreq, which is freely available at https://github.com/fanagislab/kmerfreq.

📄 PDF Abstract BibTeX arXiv:2003.11817

Code (1)

fanagislab/kmerfreq 공식 구현

Similar Papers 제목 키워드 기반

A spectral algorithm for fast de novo layout of uncorrected long nanopore reads

2017-07-17

Motivation: New long read sequencers promise to transform sequencing and genome assembly by producing reads tens of kilobases long. However their high error rate significantly complicates assembly and requires expensive …

Lerna: Transformer Architectures for Configuring Error Correction Tools for Short- and Long-Read Genome Sequencing

2021-12-19 · Atul Sharma, Pranjal Jain, Ashraf Mahgoub, Zihan Zhou 외

Sequencing technologies are prone to errors, making error correction (EC) necessary for downstream applications. EC tools need to be manually configured for optimal performance. We find that the optimal parameters (e.g.,…

GPULanguage Modelling

Inverse population genetic problems with noise: inferring extent and structure of haplotype blocks from point allele frequencies

2024-06-20 · Oliver Keatinge Clay

A haplotype block, or simply a block, is a chromosomal segment, DNA base sequence or string that occurs in only a few variants or types in the genomes of a population of interest, and that has an encapsulated or 'private…

Position

Identifying 3D Genome Organization in Diploid Organisms via Euclidean Distance Geometry

2021-01-13 · Anastasiya Belyaeva, Kaie Kubjas, Lawrence J. Sun, Caroline Uhler

The spatial organization of the DNA in the cell nucleus plays an important role for gene regulation, DNA replication, and genomic integrity. Through the development of chromosome conformation capture experiments (such as…

Correcting for Cryptic Relatedness in Genome-Wide Association Studies

2016-02-25

While the individuals chosen for a genome-wide association study (GWAS) may not be closely related to each other, there can be distant (cryptic) relationships that confound the evidence of disease association. These cryp…