Modelling Genetic Variations using Fragmentation-Coagulation Processes
We propose a novel class of Bayesian nonparametric models for sequential data called fragmentation-coagulation processes (FCPs). FCPs model a set of sequences using a partition-valued Markov process which evolves by splitting and merging clusters. An FCP is exchangeable, projective, stationary and reversible, and its equilibrium distributions are given by the Chinese restaurant process. As opposed to hidden Markov models, FCPs allow for flexible modelling of the number of clusters, and they avoid label switching non-identifiability problems. We develop an efficient Gibbs sampler for FCPs which uses uniformization and the forward-backward algorithm. Our development of FCPs is motivated by applications in population genetics, and we demonstrate the utility of FCPs on problems of genotype imputation with phased and unphased SNP data.
Code (0)
등록된 구현이 없습니다.
Tasks
ImputationSimilar Papers 제목 키워드 기반
Scalable imputation of genetic data with a discrete fragmentation-coagulation process
We present a Bayesian nonparametric model for genetic sequence data in which a set of genetic sequences is modelled using a Markov model of partitions. The partitions at consecutive locations in the genome are related b…
Computational EfficiencyImputationModelling silicosis: dynamics of a model with piecewise constant rate coefficients
We study the dynamics about equilibria of an infinite dimension coagulation-fragmentation-death model for the silicosis disease mechanism introduced recently by da Costa, Drmota, and Grinfeld [Modelling silicosis: struct…
MathFragmentation Coagulation Based Mixed Membership Stochastic Blockmodel
The Mixed-Membership Stochastic Blockmodel~(MMSB) is proposed as one of the state-of-the-art Bayesian relational methods suitable for learning the complex hidden structure underlying the network data. However, the curren…
ClusteringA dependent partition-valued process for multitask clustering and time evolving network modelling
The fundamental aim of clustering algorithms is to partition data points. We consider tasks where the discovered partition is allowed to vary with some covariate such as space or time. One approach would be to use fragme…
ClusteringGaussian ProcessesTime SeriesTime Series AnalysisImprovements to the Sequence Memoizer
The sequence memoizer is a model for sequence data with state-of-the-art performance on language modeling and compression. We propose a number of improvements to the model and inference algorithm, including an enlarged r…
Language ModelingLanguage Modelling