paper-with-me

Papers

Linear normalised hash function for clustering gene sequences and identifying reference sequences from multiple sequence alignments

2023-11-29 · Manal Helal, Fanrong Kong, Sharon C-A Chen, Fei Zhou, Dominic E Dwyer, John Potter, Vitali Sintchenko

The aim of this study was to develop a method that would identify the cluster centroids and the optimal number of clusters for a given sensitivity level and could work equally well for the different sequence datasets. A novel method that combines the linear mapping hash function and multiple sequence alignment (MSA) was developed. This method takes advantage of the already sorted by similarity sequences from the MSA output, and identifies the optimal number of clusters, clusters cut-offs, and clusters centroids that can represent reference gene vouchers for the different species. The linear mapping hash function can map an already ordered by similarity distance matrix to indices to reveal gaps in the values around which the optimal cut-offs of the different clusters can be identified. The method was evaluated using sets of closely related (16S rRNA gene sequences of Nocardia species) and highly variable (VP1 genomic region of Enterovirus 71) sequences and outperformed existing unsupervised machine learning clustering methods and dimensionality reduction methods. This method does not require prior knowledge of the number of clusters or the distance between clusters, handles clusters of different sizes and shapes, and scales linearly with the dataset. The combination of MSA with the linear mapping hash function is a computationally efficient way of gene sequence clustering and can be a valuable tool for the assessment of similarity, clustering of different microbial genomes, identifying reference sequences, and for the study of evolution of bacteria and viruses.

📄 PDF Abstract BibTeX arXiv:2311.17964

Code (0)

등록된 구현이 없습니다.

Tasks

ClusteringDimensionality ReductionMultiple Sequence Alignment

Similar Papers 제목 키워드 기반

BCTH: A Novel Text Hashing Approach via Bayesian Clustering

2020-12-01 · Asian Chapter of the Association for Computational Linguistics 2020 · Ying Wenjie, Yuquan Le, Hantao Xiong

Similarity search is to find the most similar items for a certain target item. The ability of similarity search at large scale plays a significant role in many information retrieval applications, and thus has received mu…

ClusteringInformation RetrievalRetrieval

Improving Spectral Clustering using the Asymptotic Value of the Normalised Cut

2017-03-29 · David Hofmeyr

Spectral clustering is a popular and versatile clustering method based on a relaxation of the normalised graph cut objective. Despite its popularity, however, there is no single agreed upon method for tuning the importan…

Clustering

Cluster-wise Unsupervised Hashing for Cross-Modal Similarity Search

2019-11-11 · Lu Wang, Jie Yang

Large-scale cross-modal hashing similarity retrieval has attracted more and more attention in modern search applications such as search engines and autopilot, showing great superiority in computation and storage. However…

ClusteringRetrieval

Unsupervised Deep Generative Adversarial Hashing Network

2018-06-01 · CVPR 2018 6 · Kamran Ghasedi Dizaji, Feng Zheng, Najmeh Sadoughi, Yanhua Yang 외

Unsupervised deep hash functions have not shown satisfactory improvements against the shallow alternatives, and usually, require supervised pretraining to avoid getting stuck in bad local minima. In this paper, we propos…

ClusteringImage ClusteringImage RetrievalRetrieval+1

Normalised clustering accuracy: An asymmetric external cluster validity measure

2022-09-07 · Marek Gagolewski

There is no, nor will there ever be, single best clustering algorithm. Nevertheless, we would still like to be able to distinguish between methods that work well on certain task types and those that systematically underp…

Clusteringset matching