Aligning 415 519 proteins in less than two hours on PC
Rapid development of modern sequencing platforms enabled an unprecedented growth of protein families databases. The abundance of sets composed of hundreds of thousands sequences is a great challenge for multiple sequence alignment algorithms. In the article we introduce FAMSA, a new progressive algorithm designed for fast and accurate alignment of thousands of protein sequences. Its features include the utilisation of longest common subsequence measure for determining pairwise similarities, a novel method of gap costs evaluation, and a new iterative refinement scheme. Importantly, its implementation is highly optimised and parallelised to make the most of modern computer platforms. Thanks to the above, quality indicators, namely sum-of-pairs and total-column scores, show FAMSA to be superior to competing algorithms like Clustal Omega or MAFFT for datasets exceeding a few thousand of sequences. The quality does not compromise time and memory requirements which are an order of magnitude lower than that of existing solutions. For example, a family of 415 519 sequences was analysed in less than two hours and required only 8GB of RAM. FAMSA is freely available at http://sun.aei.polsl.pl/REFRESH/famsa.
Code (0)
등록된 구현이 없습니다.
Tasks
Multiple Sequence AlignmentVocal Bursts Valence PredictionSimilar Papers 제목 키워드 기반
Variants of intrinsic disorder in the human proteome
In this paper we propose a straightforward operational definition of variants of disordered proteins, taking the human proteome as a case study. The focus is on a distinction between mostly unstructured proteins and prot…
Speech Vecalign: an Embedding-based Method for Aligning Parallel Speech Documents
We present Speech Vecalign, a parallel speech document alignment method that monotonically aligns speech segment embeddings and does not depend on text transcriptions. Compared to the baseline method Global Mining, a var…
Speech-to-Speech TranslationAn introductory guide to aligning networks using SANA, the Simulated Annealing Network Aligner
Sequence alignment has had an enormous impact on our understanding of biology, evolution, and disease. The alignment of biological {\em networks} holds similar promise. Biological networks generally model interactions be…
CPUFast Proteome Identification and Quantification from Data-Dependent Acquisition - Tandem Mass Spectrometry using Free Software Tools
Identification of nearly all proteins in a system using data-dependent acquisition (DDA) mass spectrometry has become routine for simple organisms, such as bacteria and yeast. Still, quantification of the identified prot…
Training self-supervised peptide sequence models on artificially chopped proteins
Representation learning for proteins has primarily focused on the global understanding of protein sequences regardless of their length. However, shorter proteins (known as peptides) take on distinct structures and functi…
Data AugmentationLanguage ModelingLanguage ModellingRepresentation Learning+1