A draft genome assembly of southern bluefin tuna Thunnus maccoyii
Tuna are large pelagic fish whose populations are close to panmixia. In addition, they are threatened species, so it is important for the maintenance and monitoring of genetic diversity that genetic information at a genome level be obtained. Here we report the draft assembly of the southern bluefin tuna genome and the collection of genome-wide sequence data for five other tuna species. We sampled five tuna species of the genus Thunnus, the northern and southern bluefin, yellowfin, albacore, and bigeye, as well as the skipjack (Katsuwonis pelamis), a tuna-like species. Genome assembly was facilitated at k-mer=25 while k-mer=51 generated assembly artefacts. The estimated size of the southern bluefin tuna genome was 795 Mb. We assembled two southern bluefin tuna individuals independently using both paired end and mate pair sequence. This resulted in scaffolds with N50>174,000 bp and maximum scaffold lengths>1.4 Mb. Our estimate of the size of the assembled genome was the scaffolded sequences in common to both assemblies, which amounted to 721 Mb of the 795 Mb of the southern bluefin tuna genome sequence. Using BLAST, there were matches between 13,039 of 14,341 (91%) refseq mRNA of the zebrafish Danio rerio to the tuna assembly indicating that most of a generic fish transcriptome was covered by the assembly.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
GapPredict: A Language Model for Resolving Gaps in Draft Genome Assemblies
Short-read DNA sequencing instruments can yield over 1e+12 bases per run, typically composed of reads 150 bases long. Despite this high throughput, de novo assembly algorithms have difficulty reconstructing contiguous ge…
Language ModelingLanguage ModellingOrienting Ordered Scaffolds: Complexity and Algorithms
Despite the recent progress in genome sequencing and assembly, many of the currently available assembled genomes come in a draft form. Such draft genomes consist of a large number of genomic fragments (scaffolds), whose …
Apollo: A Sequencing-Technology-Independent, Scalable, and Accurate Assembly Polishing Algorithm
Long reads produced by third-generation sequencing technologies are used to construct an assembly (i.e., the subject's genome), which is further used in downstream genome analysis. Unfortunately, long reads have high seq…
ntLink: a toolkit for de novo genome assembly scaffolding and mapping using long reads
With the increasing affordability and accessibility of genome sequencing data, de novo genome assembly is an important first step to a wide variety of downstream studies and analyses. Therefore, bioinformatics tools that…
Computational EfficiencyEstimation of genomic characteristics by analyzing k-mer frequency in de novo genome projects
Background: With the fast development of next generation sequencing technologies, increasing numbers of genomes are being de novo sequenced and assembled. However, most are in fragmental and incomplete draft status, and …