Bayesian Pedigree Analysis using Measure Factorization
Pedigrees, or family trees, are directed graphs used to identify sites of the genome that are correlated with the presence or absence of a disease. With the advent of genotyping and sequencing technologies, there has been an explosion in the amount of data available, both in the number of individuals and in the number of sites. Some pedigrees number in the thousands of individuals. Meanwhile, analysis methods have remained limited to pedigrees of <100 individuals which limits analyses to many small independent pedigrees. Disease models, such those used for the linkage analysis log-odds (LOD) estimator, have similarly been limited. This is because linkage anlysis was originally designed with a different task in mind, that of ordering the sites in the genome, before there were technologies that could reveal the order. LODs are difficult to interpret and nontrivial to extend to consider interactions among sites. These developments and difficulties call for the creation of modern methods of pedigree analysis. Drawing from recent advances in graphical model inference and transducer theory, we introduce a simple yet powerful formalism for expressing genetic disease models. We show that these disease models can be turned into accurate and efficient estimators. The technique we use for constructing the variational approximation has potential applications to inference in other large-scale graphical models. This method allows inference on larger pedigrees than previously analyzed in the literature, which improves disease site prediction.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Non-Identifiable Pedigrees and a Bayesian Solution
Some methods aim to correct or test for relationships or to reconstruct the pedigree, or family tree. We show that these methods cannot resolve ties for correct relationships due to identifiability of the pedigree likeli…
Efficient Reconstruction of Stochastic Pedigrees
We introduce a new algorithm called {\sc Rec-Gen} for reconstructing the genealogy or \textit{pedigree} of an extant population purely from its genetic data. We justify our approach by giving a mathematical proof of the …
Blang: Bayesian declarative modelling of general data structures and inference via algorithms based on distribution continua
Consider a Bayesian inference problem where a variable of interest does not take values in a Euclidean space. These "non-standard" data structures are in reality fairly common. They are frequently used in problems involv…
Bayesian InferenceAn nth-cousin mating model and the n-anacci numbers
In seeking to understand the size of inbred pedigrees, J. Lachance (J. Theor. Biol. 261, 238-247, 2009) studied a population model in which, for a fixed value of $n$, each mating occurs between $n$th cousins. We explain …
Conditional gene genealogies given the population pedigree for a diploid Moran model with selfing
We introduce a stochastic model of a population with overlapping generations and arbitrary levels of self-fertilization versus outcrossing. We study how the global graph of reproductive relationships, or population pedig…