paper-with-me

홈 › Papers

On the quartet distance given partial information

2021-11-25 · Sagi Snir, Osnat Weissberg, Raphael Yuster

Let $T$ be an arbitrary phylogenetic tree with $n$ leaves. It is well-known that the average quartet distance between two assignments of taxa to the leaves of $T$ is $\frac 23 \binom{n}{4}$. However, a longstanding conjecture of Bandelt and Dress asserts that $(\frac 23 +o(1))\binom{n}{4}$ is also the {\em maximum} quartet distance between two assignments. While Alon, Naves, and Sudakov have shown this indeed holds for caterpillar trees, the general case of the conjecture is still unresolved. A natural extension is when partial information is given: the two assignments are known to coincide on a given subset of taxa. The partial information setting is biologically relevant as the location of some taxa (species) in the phylogenetic tree may be known, and for other taxa it might not be known. What can we then say about the average and maximum quartet distance in this more general setting? Surprisingly, even determining the {\em average} quartet distance becomes a nontrivial task in the partial information setting and determining the maximum quartet distance is even more challenging, as these turn out to be dependent of the structure of $T$. In this paper we prove nontrivial asymptotic bounds that are sometimes tight for the average quartet distance in the partial information setting. We also show that the Bandelt and Dress conjecture does not generally hold under the partial information setting. Specifically, we prove that there are cases where the average and maximum quartet distance substantially differ.

📄 PDF Abstract BibTeX arXiv:2111.13101

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Inferring metric trees from weighted quartets via an intertaxon distance

2020-02-11

A metric phylogenetic tree relating a collection of taxa induces weighted rooted triples and weighted quartets for all subsets of three and four taxa, respectively. New intertaxon distances are defined that can be calcul…

Automating Sound Change Prediction for Phylogenetic Inference: A Tukanoan Case Study

2024-02-02 · Kalvin Chang, Nathaniel R. Robinson, Anna Cai, Ting Chen 외

We describe a set of new methods to partially automate linguistic phylogenetic inference given (1) cognate sets with their respective protoforms and sound laws, (2) a mapping from phones to their articulatory features an…

A Fast Quartet Tree Heuristic for Hierarchical Clustering

2014-09-12 · Rudi L. Cilibrasi, Paul M. B. Vitanyi

The Minimum Quartet Tree Cost problem is to construct an optimal weight tree from the $3{n \choose 4}$ weighted quartet topologies on $n$ objects, where optimality means that the summed weight of the embedded quartet top…

Clusteringglobal-optimization

Towards identifying the optimal datasize for lexically-based Bayesian inference of linguistic phylogenies

2018-08-01 · COLING 2018 8 · Taraka Rama, S{\o}ren Wichmann

Bayesian linguistic phylogenies are standardly based on cognate matrices for words referring to a fix set of meanings{---}typically around 100-200. To this day there has not been any empirical investigation into which da…

Bayesian Inference

Designing weights for quartet-based methods when data is heterogeneous across lineages

2022-02-27 · Marta Casanellas, Jesús Fernández-Sánchez, Marina Garrote-López, Marc Sabaté-Vidales

Homogeneity across lineages is a common assumption in phylogenetics according to which nucleotide substitution rates remain constant in time and do not depend on lineages. This is a simplifying hypothesis which is often …