Inverse population genetic problems with noise: inferring extent and structure of haplotype blocks from point allele frequencies
A haplotype block, or simply a block, is a chromosomal segment, DNA base sequence or string that occurs in only a few variants or types in the genomes of a population of interest, and that has an encapsulated or 'private' frequency distribution of the string types that is not shared by neighbouring blocks or regions on the same chromosome. We consider two inverse problems of genetic interest: from just the frequencies of the symbol types (4 base types, possible single-base alleles) at each position (point, base/nucleotide) along the string, infer the location of the left and right boundaries of the block (block extent), and the number and relative frequencies of the string types occurring in the block (block structure). The large majority of variable positions in human and also other (e.g., fungal) genomes appear to be biallelic, i.e., the position allows only a choice between two possible symbols. The symbols can then be encoded as 0 (major) and 1 (minor), or as $\uparrow$ and $\downarrow$ as in Ising models, so the scenario reduces to problems on Boolean strings/bitstrings and Boolean matrices. The specifying of major allele frequencies (MAF) as used in genetics fits naturally into this framework. A simple example from human chromosome 9 is presented.
Code (0)
등록된 구현이 없습니다.
Tasks
PositionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Connecting the Dots: Range Expansions across Landscapes with Quenched Noise
When biological populations expand into new territory, the evolutionary outcomes can be strongly influenced by genetic drift, the random fluctuations in allele frequencies. Meanwhile, spatial variability in the environme…
Graph Learning for Inverse Landscape Genetics
The problem of inferring unknown graph edges from numerical data at a graph's nodes appears in many forms across machine learning. We study a version of this problem that arises in the field of \emph{landscape genetics},…
Graph LearningLarger Offspring Populations Help the $(1 + (λ, λ))$ Genetic Algorithm to Overcome the Noise
Evolutionary algorithms are known to be robust to noise in the evaluation of the fitness. In particular, larger offspring population sizes often lead to strong robustness. We analyze to what extent the $(1+(\lambda,\lamb…
Evolutionary AlgorithmsBayesian hierarchical modelling for inferring genetic interactions in yeast
Quantitative Fitness Analysis (QFA) is a high-throughput experimental and computational methodology for measuring the growth of microbial populations. QFA screens can be used to compare the health of cell populations wit…
Experimental DesignNovel probabilistic models of spatial genetic ancestry with applications to stratification correction in genome-wide association studies
Genetic variation in human populations is influenced by geographic ancestry due to spatial locality in historical mating and migration patterns. Spatial population structure in genetic datasets has been traditionally ana…