Towards more realistic models of genomes in populations: the Markov-modulated sequentially Markov coalescent
The development of coalescent theory paved the way to statistical inference from population genetic data. In the genomic era, however, coalescent models are limited due to the complexity of the underlying ancestral recombination graph. The sequentially Markov coalescent (SMC) is a heuristic that enables the modelling of complete genomes under the coalescent framework. While it empowers the inference of detailed demographic history of a population from as few as one diploid genome, current implementations of the SMC make unrealistic assumptions about the homogeneity of the coalescent process along the genome, ignoring the intrinsic spatial variability of parameters such as the recombination rate. Here, I review the historical developments of SMC models and discuss the evidence for parameter heterogeneity. I then survey approaches to handle this heterogeneity, focusing on a recently developed extension of the SMC.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Sex as Gibbs Sampling: a probability model of evolution
We show that evolutionary computation can be implemented as standard Markov-chain Monte-Carlo (MCMC) sampling. With some care, `genetic algorithms' can be constructed that are reversible Markov chains that satisfy detail…
Privacy-hardened and hallucination-resistant synthetic data generation with logic-solvers
Machine-generated data is a valuable resource for training Artificial Intelligence algorithms, evaluating rare workflows, and sharing data under stricter data legislations. The challenge is to generate data that is accur…
Generative Adversarial NetworkHallucinationSynthetic Data GenerationKGP: An R Package with Metadata from the 1000 Genomes Project
The 1000 Genomes Project provides sequencing data on 3,202 samples from 26 populations spanning five continental regions with no access or use restrictions. The kgp R package provides consistent and comprehensive metadat…
Application of Markov Structure of Genomes to Outlier Identification and Read Classification
In this paper we apply the structure of genomes as second-order Markov processes specified by the distributions of successive triplets of bases to two bioinformatics problems: identification of outliers in genome databas…
CLMB: deep contrastive learning for robust metagenomic binning
The reconstruction of microbial genomes from large metagenomic datasets is a critical procedure for finding uncultivated microbial populations and defining their microbial functional roles. To achieve that, we need to pe…
BenchmarkingContrastive LearningDenoising