paper-with-me

홈 › Papers

k-mer-based approaches to bridging pangenomics and population genetics

2024-09-18 · Miles D. Roberts, Olivia Davis, Emily B. Josephs, Robert J. Williamson

Many commonly studied species now have more than one chromosome-scale genome assembly, revealing a large amount of genetic diversity previously missed by approaches that map short reads to a single reference. However, many species still lack multiple reference genomes and correctly aligning references to build pangenomes is challenging, limiting our ability to study this missing genomic variation in population genetics. Here, we argue that $k$-mers are a crucial stepping stone to bridging the reference-focused paradigms of population genetics with the reference-free paradigms of pangenomics. We review current literature on the uses of $k$-mers for performing three core components of most population genetics analyses: identifying, measuring, and explaining patterns of genetic variation. We also demonstrate how different $k$-mer-based measures of genetic variation behave in population genetic simulations according to the choice of $k$, depth of sequencing coverage, and degree of data compression. Overall, we find that $k$-mer-based measures of genetic diversity scale consistently with pairwise nucleotide diversity ($\pi$) up to values of about $\pi = 0.025$ ($R^2 = 0.97$) for neutrally evolving populations. For populations with even more variation, using shorter $k$-mers will maintain the scalability up to at least $\pi = 0.1$. Furthermore, in our simulated populations, $k$-mer dissimilarity values can be reliably approximated from counting bloom filters, highlighting a potential avenue to decreasing the memory burden of $k$-mer based genomic dissimilarity analyses. For future studies, there is a great opportunity to further develop methods to identifying selected loci using $k$-mers.

📄 PDF Abstract BibTeX arXiv:2409.11683

Code (0)

등록된 구현이 없습니다.

Tasks

Data CompressionDiversity

Methods 이 논문이 사용한 방법론

BLOOM BLOOM is a decoder-only Transformer language model that was trained on the ROOTS corpus, a dataset comprising hundreds of sources in 46 natural and 13 programming languages…

Similar Papers 제목 키워드 기반

Bridging Wright-Fisher and Moran models

2024-07-17 · Arthur Alexandre, Alia Abbara, Cecilia Fruet, Claude Loverdo 외

The Wright-Fisher model and the Moran model are both widely used in population genetics. They describe the time evolution of the frequency of an allele in a well-mixed population with fixed size. We propose a simple and …

Population Genetics with Fluctuating Population Sizes

2016-08-29

Standard neutral population genetics theory with a strictly fixed population size has important limitations. An alternative model that allows independently fluctuating population sizes and reproduces the standard neutral…

Population genetics: an introduction for physicists

2024-08-05 · Andrea Iglesias-Ramas, Samuele Pio Lipani, Rosalind J. Allen

Population genetics lies at the heart of evolutionary theory. This topic forms part of many biological science curricula but is rarely taught to physics students. Since physicists are becoming increasingly interested in …

Population Genetics and Evolution

2018-03-22

These lecture notes introduce key concepts of mathematical population genetics within the most elementary setting and describe a few recent applications to microbial evolution experiments. Pointers to the literature for …

Human genetic admixture through the lens of population genomics

2021-09-24 · Shyamalika Gopalan, Samuel Patillo Smith, Katharine Korunes, Iman Hamid 외

Over the last fifty years, geneticists have made great strides in understanding how our species' evolutionary history gave rise to current patterns of human genetic diversity classically summarized by Lewontin in his 197…

Diversity