paper-with-me

홈 › Papers

Cnidaria: fast, reference-free clustering of raw and assembled genome and transcriptome NGS data

2015-11-17

Background: Identification of biological specimens is a major requirement for a range of applications. Reference-free methods analyse unprocessed sequencing data without relying on prior knowledge, but generally do not scale to arbitrarily large genomes and arbitrarily large phylogenetic distances. Results: We present Cnidaria, a practical tool for clustering genomic and transcriptomic data with no limitation on genome size or phylogenetic distances. We successfully simultaneously clustered 169 genomic and transcriptomic datasets from 4 kingdoms, achieving 100% identification accuracy at supra-species level and 78% accuracy for species level. Discussion: CNIDARIA allows for fast, resource-efficient comparison and identification of both raw and assembled genome and transcriptome data. This can help answer both fundamental (e.g. in phylogeny, ecological diversity analysis) and practical questions (e.g. sequencing quality control, primer design).

📄 PDF Abstract BibTeX arXiv:1511.05530

Code (0)

등록된 구현이 없습니다.

Tasks

ClusteringDiversity

Similar Papers 제목 키워드 기반

Reads2Vec: Efficient Embedding of Raw High-Throughput Sequencing Reads Data

2022-11-15 · Prakash Chourasia, Sarwan Ali, Simone Ciccolella, Gianluca Della Vedova 외

The massive amount of genomic data appearing for SARS-CoV-2 since the beginning of the COVID-19 pandemic has challenged traditional methods for studying its dynamics. As a result, new methods such as Pangolin, which can …

ClusteringVocal Bursts Intensity Prediction

Multiple Kernel Clustering with Dual Noise Minimization

2022-07-13 · Junpu Zhang, Liang Li, Siwei Wang, Jiyuan Liu 외

Clustering is a representative unsupervised method widely applied in multi-modal and multi-view scenarios. Multiple kernel clustering (MKC) aims to group data by integrating complementary information from base kernels. A…

Clustering

$DC^2$: A Divide-and-conquer Algorithm for Large-scale Kernel Learning with Application to Clustering

2019-11-16 · Ke Alexander Wang, Xinran Bian, Pan Liu, Donghui Yan

Divide-and-conquer is a general strategy to deal with large scale problems. It is typically applied to generate ensemble instances, which potentially limits the problem size it can handle. Additionally, the data are ofte…

Clustering

Towards complete representation of bacterial contents in metagenomic samples

2022-09-30 · Xiaowen Feng, Heng Li

Background: In the metagenome assembly of a microbiome community, we may think abundant species would be easier to assemble due to their deeper coverage. However, this conjucture is rarely tested. We often do not know ho…

Diversity

ViralVectors: Compact and Scalable Alignment-free Virome Feature Generation

2023-04-06 · Sarwan Ali, Prakash Chourasia, Zahra Tayebi, Babatunde Bello 외

The amount of sequencing data for SARS-CoV-2 is several orders of magnitude larger than any virus. This will continue to grow geometrically for SARS-CoV-2, and other viruses, as many countries heavily finance genomic sur…

4kDecision Making