Limits on Inferring T-cell Specificity from Partial Information
A key challenge in molecular biology is to decipher the mapping of protein sequence to function. To perform this mapping requires the identification of sequence features most informative about function. Here, we quantify the amount of information (in bits) that T-cell receptor (TCR) sequence features provide about antigen specificity. We identify informative features by their degree of conservation among antigen-specific receptors relative to null expectations. We find that TCR specificity synergistically depends on the hypervariable regions of both receptor chains, with a degree of synergy that strongly depends on the ligand. Using a coincidence-based approach to measuring information enables us to directly bound the accuracy with which TCR specificity can be predicted from partial matches to reference sequences. We anticipate that our statistical framework will be of use for developing machine learning models for TCR specificity prediction and for optimizing TCRs for cell therapies. The proposed coincidence-based information measures might find further applications in bounding the performance of pairwise classifiers in other fields.
Code (0)
등록된 구현이 없습니다.
Tasks
SpecificitySimilar Papers 제목 키워드 기반
Attention-aware contrastive learning for predicting T cell receptor-antigen binding specificity
It has been verified that only a small fraction of the neoantigens presented by MHC class I molecules on the cell surface can elicit T cells. The limitation can be attributed to the binding specificity of T cell receptor…
Contrastive LearningSpecificityPleiotropy enables specific and accurate signaling in the presence of ligand cross talk
Living cells sense their environment through the binding of extra-cellular molecular ligands to cell surface receptors. Puzzlingly, vast numbers of signaling pathways exhibit a high degree of cross talk between different…
SpecificityLearning Agile Skills via Adversarial Imitation of Rough Partial Demonstrations
Learning agile skills is one of the main challenges in robotics. To this end, reinforcement learning approaches have achieved impressive results. These methods require explicit task information in terms of a reward funct…
Reinforcement Learning (RL)Encoding Domain Information with Sparse Priors for Inferring Explainable Latent Variables
Latent variable models are powerful statistical tools that can uncover relevant variation between patients or cells, by inferring unobserved hidden states from observable high-dimensional data. A major shortcoming of cur…
Representation of ambiguity in pretrained models and the problem of domain specificity
Recent developments in pretrained language models have led to many advances in NLP. These models have excelled at learning powerful contextual representations from very large corpora. Fine-tuning these models for downstr…
Specificity