Influence of Multiple Sequence Alignment Depth on Potts Statistical Models of Protein Covariation
Potts statistical models have become a popular and promising way to analyze mutational covariation in protein Multiple Sequence Alignments (MSAs) in order to understand protein structure, function and fitness. But the statistical limitations of these models, which can have millions of parameters and are fit to MSAs of only thousands or hundreds of effective sequences using a procedure known as inverse Ising inference, are incompletely understood. In this work we predict how model quality degrades as a function of the number of sequences $N$, sequence length $L$, amino-acid alphabet size $q$, and the degree of conservation of the MSA, in different applications of the Potts models: In "fitness" predictions of individual protein sequences, in predictions of the effects of single-point mutations, in "double mutant cycle" predictions of epistasis, and in 3-d contact prediction in protein structure. We show how as MSA depth $N$ decreases an "overfitting" effect occurs such that sequences in the training MSA have overestimated fitness, and we predict the magnitude of this effect and discuss how regularization can help correct for it, use a regularization procedure motivated by statistical analysis of the effects of finite sampling. We find that as $N$ decreases the quality of point-mutation effect predictions degrade least, fitness and epistasis predictions degrade more rapidly, and contact predictions are most affected. However, overfitting becomes negligible for MSA depths of more than a few thousand effective sequences, as often used in practice, and regularization becomes less necessary. We discuss the implications of these results for users of Potts covariation analysis.
Code (0)
등록된 구현이 없습니다.
Tasks
Multiple Sequence AlignmentSimilar Papers 제목 키워드 기반
Neural Potts Model
We propose the Neural Potts Model objective as an amortized optimization problem. The objective enables training a single model with shared parameters to explicitly model energy landscapes across multiple protein familie…
modelGenerative power of a protein language model trained on multiple sequence alignments
Computational models starting from large ensembles of evolutionarily related protein sequences capture a representation of protein families and learn constraints associated to protein structure and function. They thus op…
Language ModelingLanguage ModellingMasked Language ModelingProtein Design+1Improving contact prediction along three dimensions
Correlation patterns in multiple sequence alignments of homologous proteins can be exploited to infer information on the three-dimensional structure of their members. The typical pipeline to address this task, which we i…
PredictionBenchmarking inverse statistical approaches for protein structure and design with exactly solvable models
Inverse statistical approaches to determine protein structure and function from Multiple Sequence Alignments (MSA) are emerging as powerful tools in computational biology. However the underlying assumptions of the relati…
BenchmarkingProtein language models trained on multiple sequence alignments learn phylogenetic relationships
Self-supervised neural language models with attention have recently been applied to biological sequence data, advancing structure, function and mutational effect prediction. Some protein language models, including MSA Tr…
Prediction