paper-with-me

홈 › Papers

ODBO: Bayesian Optimization with Search Space Prescreening for Directed Protein Evolution

2022-05-19 · Lixue Cheng, ZiYi Yang, ChangYu Hsieh, Benben Liao, Shengyu Zhang

Directed evolution is a versatile technique in protein engineering that mimics the process of natural selection by iteratively alternating between mutagenesis and screening in order to search for sequences that optimize a given property of interest, such as catalytic activity and binding affinity to a specified target. However, the space of possible proteins is too large to search exhaustively in the laboratory, and functional proteins are scarce in the vast sequence space. Machine learning (ML) approaches can accelerate directed evolution by learning to map protein sequences to functions without building a detailed model of the underlying physics, chemistry and biological pathways. Despite the great potentials held by these ML methods, they encounter severe challenges in identifying the most suitable sequences for a targeted function. These failures can be attributed to the common practice of adopting a high-dimensional feature representation for protein sequences and inefficient search methods. To address these issues, we propose an efficient, experimental design-oriented closed-loop optimization framework for protein directed evolution, termed ODBO, which employs a combination of novel low-dimensional protein encoding strategy and Bayesian optimization enhanced with search space prescreening via outlier detection. We further design an initial sample selection strategy to minimize the number of experimental samples for training ML models. We conduct and report four protein directed evolution experiments that substantiate the capability of the proposed framework for finding of the variants with properties of interest. We expect the ODBO framework to greatly reduce the experimental cost and time cost of directed evolution, and can be further generalized as a powerful tool for adaptive experimental design in a broader context.

📄 PDF Abstract BibTeX arXiv:2205.09548

Code (1)

sherrylixuecheng/odbo 공식 구현 pytorch

Tasks

Bayesian OptimizationExperimental DesignOutlier Detection

Similar Papers 제목 키워드 기반

TWEET-FID: An Annotated Dataset for Multiple Foodborne Illness Detection Tasks

2022-05-22 · LREC 2022 6 · Ruofan Hu, Dongyu Zhang, Dandan Tao, Thomas Hartvigsen 외

Foodborne illness is a serious but preventable public health problem -- with delays in detecting the associated outbreaks resulting in productivity loss, expensive recalls, public safety hazards, and even loss of life. W…

slot-fillingSlot Filling

Machine-learned epidemiology: real-time detection of foodborne illness at scale

2018-12-05 · Adam Sadilek, Stephanie Caty, Lauren DiPrete, Raed Mansour 외

Machine learning has become an increasingly powerful tool for solving complex problems, and its application in public health has been underutilized. The objective of this study is to test the efficacy of a machine-learne…

Epidemiology

Seasonality Patterns in 311-Reported Foodborne Illness Cases and Machine Learning-Identified Indications of Foodborne Illnesses from Yelp Reviews, New York City, 2022-2023

2024-05-09 · Eden Shaveet, Crystal Su, Daniel Hsu, Luis Gravano

Restaurants are critical venues at which to investigate foodborne illness outbreaks due to shared sourcing, preparation, and distribution of foods. Formal channels to report illness after food consumption, such as 311, N…

UCE-FID: Using Large Unlabeled, Medium Crowdsourced-Labeled, and Small Expert-Labeled Tweets for Foodborne Illness Detection

2023-12-02 · Ruofan Hu, Dongyu Zhang, Dandan Tao, Huayi Zhang 외

Foodborne illnesses significantly impact public health. Deep learning surveillance applications using social media data aim to detect early warning signals. However, labeling foodborne illness-related tweets for model tr…

Boosting SISSO Performance on Small Sample Datasets by Using Random Forests Prescreening for Complex Feature Selection

2024-09-28 · Xiaolin Jiang, Guanqi Liu, Jiaying Xie, Zhenpeng Hu

In materials science, data-driven methods accelerate material discovery and optimization while reducing costs and improving success rates. Symbolic regression is a key to extracting material descriptors from large datase…

feature selectionregressionSymbolic Regression