Accelerating high-throughput virtual screening through molecular pool-based active learning
Structure-based virtual screening is an important tool in early stage drug discovery that scores the interactions between a target protein and candidate ligands. As virtual libraries continue to grow (in excess of $10^8$ molecules), so too do the resources necessary to conduct exhaustive virtual screening campaigns on these libraries. However, Bayesian optimization techniques can aid in their exploration: a surrogate structure-property relationship model trained on the predicted affinities of a subset of the library can be applied to the remaining library members, allowing the least promising compounds to be excluded from evaluation. In this study, we assess various surrogate model architectures, acquisition functions, and acquisition batch sizes as applied to several protein-ligand docking datasets and observe significant reductions in computational costs, even when using a greedy acquisition strategy; for example, 87.9% of the top-50000 ligands can be found after testing only 2.4% of a 100M member library. Such model-guided searches mitigate the increasing computational costs of screening increasingly large virtual libraries and can accelerate high-throughput virtual screening campaigns with applications beyond docking.
Code (1)
Tasks
Active LearningBayesian OptimizationDrug DiscoveryVocal Bursts Intensity PredictionSimilar Papers 제목 키워드 기반
Docking-based Virtual Screening with Multi-Task Learning
Machine learning shows great potential in virtual screening for drug discovery. Current efforts on accelerating docking-based virtual screening do not consider using existing data of other previously developed targets. T…
BIG-bench Machine LearningDrug DiscoveryMulti-Task LearningOptimal Decision Making in High-Throughput Virtual Screening Pipelines
The need for efficient computational screening of molecular candidates that possess desired properties frequently arises in various scientific and engineering problems, including drug discovery and materials design. Howe…
Decision MakingDrug DiscoveryProperty PredictionVocal Bursts Intensity PredictionMitigating Molecular Aggregation in Drug Discovery with Predictive Insights from Explainable AI
Herein, we present the application of MEGAN, our explainable AI (xAI) model, for the identification of small colloidally aggregating molecules (SCAMs). This work offers solutions to the long-standing problem of false pos…
Drug DiscoveryTransfer learning discovery of molecular modulators for perovskite solar cells
The discovery of effective molecular modulators is essential for advancing perovskite solar cells (PSCs), but the research process is hindered by the vastness of chemical space and the time-consuming and expensive trial-…
Transfer LearningHigh Throughput Virtual Screening with Data Level Parallelism in Multi-core Processors
Improving the throughput of molecular docking, a computationally intensive phase of the virtual screening process, is a highly sought area of research since it has a significant weight in the drug designing process. With…
Molecular Docking