Orthologs from maxmer sequence context
Context-dependent identification of orthologs customarily relies on conserved gene order or whole-genome sequence alignment. It is shown here that short-range context--as short as single maximal matches--also provides an effective means to identify orthologs within whole genomes. On pristine (un-repeatmasked) mammalian whole-genome assemblies we perform a genome "intersection" that in general consumes less than one thirtieth of the computation time required by commonly used methods for whole-genome alignment, and we extract "non-embedded maximal matches," maximal matches that are not embedded into other maximal matches, as potential orthologs. An ortholog identified via non-embedded maximal matches is analogous to a "positional ortholog" or a "primary ortholog" as defined in previous literature; such orthologs constitute homologs derived from the same direct ancestor whose ancestral positions in the genome are conserved. At the nucleotide level, non-embedded maximal matches recapitulate most exact matches identified by a Lastz net alignment. At the gene level, reciprocal best hits of genes containing non-embedded maximal matches recover one-to-one orthologs annotated by Ensembl Compara with high selectivity and high sensitivity; these reciprocal best hits additionally include putatively novel orthologs not found in Ensembl (e.g. over two thousand for human/chimpanzee). The method is especially suitable for genome-wide identification of orthologs.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Construction of Gene and Species Trees from Sequence Data incl. Orthologs, Paralogs, and Xenologs
Phylogenetic reconstruction aims at finding plausible hypotheses of the evolutionary history of genes or species based on genomic sequence information. The distinction of orthologous genes (genes that having a common anc…
Leveraging Protein Language Model Embeddings for Catalytic Turnover Prediction of Adenylate Kinase Orthologs in a Low-Data Regime
Accurate prediction of enzymatic activity from amino acid sequences could drastically accelerate enzyme engineering for applications such as bioremediation and therapeutics development. In recent years, Protein Language …
Language ModelingLanguage ModellingPredictionProtein Language ModelSARS-CoV-2 orthologs of pathogenesis-involved small viral RNAs of SARS-CoV
Background: The COVID-19 pandemic clock is ticking and the survival of many of mankind's modern institutions and or survival of many individuals is at stake. There is a need for treatments to significantly reduce the mor…
BUSCO update: novel and streamlined workflows along with broader and deeper phylogenetic coverage for scoring of eukaryotic, prokaryotic, and viral genomes
Methods for evaluating the quality of genomic and metagenomic data are essential to aid genome assembly and to correctly interpret the results of subsequent analyses. BUSCO estimates the completeness and redundancy of pr…
Homo-RAG: Homology-Guided Retrieval-Augmented Generation for Cross-Species Gene Function Prediction
The functional annotation of genes in non-model organisms remains a significant challenge in computational biology, with 20-70% of sequenced genes lacking characterized functions. Traditional homology-based methods are o…