paper-with-me

홈 › Papers

Multiple Sequence Alignment is not a Solved Problem

2018-08-23

Multiple sequence alignment is a basic procedure in molecular biology, and it is often treated as being essentially a solved computational problem. However, this is not so, and here I review the evidence for this claim, and outline the requirements for a solution. The goal of alignment is often stated to be to juxtapose nucleotides (or their derivatives, such as amino acids) that have been inherited from a common ancestral nucleotide (although other goals are also possible). Unfortunately, this is not an operational definition, because homology (in this sense) refers to unique and unobservable historical events, and so there can be no objective mathematical function to optimize. Consequently, almost all algorithms developed for multiple sequence alignment are based on optimizing some sort of compositional similarity (similarity = homology + analogy). As a result, many, if not most, practitioners either manually modify computer-produced alignments or they perform de novo manual alignment, especially in the field of phylogenetics. So, if homology is the goal, then multiple sequence alignment is not yet a solved computational problem. Several criteria have been developed by biologists to help them identify potential homologies (compositional, ontogenetic, topographical and functional similarity, plus conjunction and congruence), and these criteria can be applied to molecular data, in principle. Current computer programs do implement one (or occasionally two) of these criteria, but no program implements them all. What is needed is a program that evaluates all of the evidence for the sequence homologies, optimizes their combination, and thus produces the best hypotheses of homology. This is basically an inference problem not an optimization problem.

📄 PDF Abstract BibTeX arXiv:1808.07717

Code (0)

등록된 구현이 없습니다.

Tasks

Multiple Sequence Alignment

Similar Papers 제목 키워드 기반

Neural Time Warping For Multiple Sequence Alignment

2020-06-29 · Keisuke Kawano, Takuro Kutsuna, Satoshi Koide

Multiple sequences alignment (MSA) is a traditional and challenging task for time-series analyses. The MSA problem is formulated as a discrete optimization problem and is typically solved by dynamic programming. However,…

Multiple Sequence AlignmentTime SeriesTime Series Analysis

Connecting Visual Experiences using Max-flow Network with Application to Visual Localization

2018-08-01 · A. H. Abdul Hafez, Nakul Agarwal, C. V. Jawahar

We are motivated by the fact that multiple representations of the environment are required to stand for the changes in appearance with time and for changes that appear in a cyclic manner. These changes are, for example, …

Autonomous NavigationVisual Localization

Bridging Sequence-Structure Alignment in RNA Foundation Models

2024-07-15 · Heng Yang, Renzhi Chen, Ke Li

The alignment between RNA sequences and structures in foundation models (FMs) has yet to be thoroughly investigated. Existing FMs have struggled to establish sequence-structure alignment, hindering the free flow of genom…

Multiple sequence alignment for short sequences

2015-11-15

Multiple sequence alignment (MSA) has been one of the most important problems in bioinformatics for more decades and it is still heavily examined by many mathematicians and biologists. However, mostly because of the prac…

Multiple Sequence Alignment

Handling tree-structured text: parsing directory pages

2021-11-24 · Sarang Shrivastava, Afreen Shaikh, Shivani Shrivastava, Chung Ming Ho 외

The determination of the reading sequence of text is fundamental to document understanding. This problem is easily solved in pages where the text is organized into a sequence of lines and vertical alignment runs the heig…

document understanding