Anomaly zones for uniformly sampled gene trees under the gene duplication and loss model
Recently, there has been interest in extending long-known results about the multispecies coalescent tree to other models of gene trees. Results about the gene duplication and loss (GDL) tree have mathematical proofs, including species tree identifiability, estimability, and sample complexity of popular algorithms like ASTRAL. Here, this work is continued by characterizing the anomaly zones of uniformly sampled gene trees. The anomaly zone for species trees is the set of parameters where some discordant gene tree occurs with the maximal probability. The detection of anomalous gene trees is an important problem in phylogenomics, as their presence renders effective estimation methods to being positively misleading. Under the multispecies coalescent, anomaly zones are known to exist for rooted species trees with as few as four species. The gene duplication and loss process is a generalization of the generalized linear-birth death process to the rooted species tree, where each edge is treated as a single timeline with exponential-rate duplication and loss. The methods and results come from a detailed probabilistic analysis of trajectories observed from this stochastic process. It is shown that anomaly zones do not exist for rooted GDL balanced trees on four species, but do exist for rooted caterpillar trees, as with the multispecies coalescent.
Code (0)
등록된 구현이 없습니다.
Tasks
Mathematical ProofsSimilar Papers 제목 키워드 기반
Sackin Indices for Labeled and Unlabeled Classes of Galled Trees
The Sackin index is an important measure for the balance of phylogenetic trees. We investigate two extensions of the Sackin index to the class of galled trees and two of its subclasses (simplex galled trees and normal ga…
Machine Learning-Based Localization Accuracy of RFID Sensor Networks via RSSI Decision Trees and CAD Modeling for Defense Applications
Radio Frequency Identification (RFID) tracking may be a viable solution for defense assets that must be stored in accordance with security guidelines. However, poor sensor specificity (vulnerabilities include long range …
Anomaly DetectionAnomaly Detection for Sparse and Irregular Multivariate Time Series with Latent SDEs
Multivariate time series anomaly detection (MTSAD) is critical for a wide range of application areas, such as industrial monitoring, cybersecurity, or healthcare. Real-world data is often sparse, irregularly sampled or p…
Time Series Anomaly DetectionNew characterizations of minimum spanning trees and of saliency maps based on quasi-flat zones
We study three representations of hierarchies of partitions: dendrograms (direct representations), saliency maps, and minimum spanning trees. We provide a new bijection between saliency maps and hierarchies based on quas…
Dynamic Decision Boundary for One-class Classifiers applied to non-uniformly Sampled Data
A typical issue in Pattern Recognition is the non-uniformly sampled data, which modifies the general performance and capability of machine learning algorithms to make accurate predictions. Generally, the data is consider…
One-class classifier