paper-with-me

홈 › Papers

Deep Learning for Virtual Screening: Five Reasons to Use ROC Cost Functions

2020-06-25 · Vladimir Golkov, Alexander Becker, Daniel T. Plop, Daniel Čuturilo, Neda Davoudi, Jeffrey Mendenhall, Rocco Moretti, Jens Meiler, Daniel Cremers

Computer-aided drug discovery is an essential component of modern drug development. Therein, deep learning has become an important tool for rapid screening of billions of molecules in silico for potential hits containing desired chemical features. Despite its importance, substantial challenges persist in training these models, such as severe class imbalance, high decision thresholds, and lack of ground truth labels in some datasets. In this work we argue in favor of directly optimizing the receiver operating characteristic (ROC) in such cases, due to its robustness to class imbalance, its ability to compromise over different decision thresholds, certain freedom to influence the relative weights in this compromise, fidelity to typical benchmarking measures, and equivalence to positive/unlabeled learning. We also propose new training schemes (coherent mini-batch arrangement, and usage of out-of-batch samples) for cost functions based on the ROC, as well as a cost function based on the logAUC metric that facilitates early enrichment (i.e. improves performance at high decision thresholds, as often desired when synthesizing predicted hit compounds). We demonstrate that these approaches outperform standard deep learning approaches on a series of PubChem high-throughput screening datasets that represent realistic and diverse drug discovery campaigns on major drug target families.

📄 PDF Abstract BibTeX arXiv:2007.07029

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingDrug Discovery

Similar Papers 제목 키워드 기반

Accelerating high-throughput virtual screening through molecular pool-based active learning

2020-12-13 · David E. Graff, Eugene I. Shakhnovich, Connor W. Coley

Structure-based virtual screening is an important tool in early stage drug discovery that scores the interactions between a target protein and candidate ligands. As virtual libraries continue to grow (in excess of $10^8$…

Active LearningBayesian OptimizationDrug DiscoveryVocal Bursts Intensity Prediction

Transfer Learning across Different Chemical Domains: Virtual Screening of Organic Materials with Deep Learning Models Pretrained on Small Molecule and Chemical Reaction Data

2023-11-30 · Chengwei Zhang, Yushuang Zhai, Ziyang Gong, Hongliang Duan 외

Machine learning is becoming a preferred method for the virtual screening of organic materials due to its cost-effectiveness over traditional computationally demanding techniques. However, the scarcity of labeled data fo…

Transfer Learning

APEX: Approximate-but-exhaustive search for ultra-large combinatorial synthesis libraries

2025-10-28 · Aryan Pedawi, Jordi Silvestre-Ryan, Bradley Worley, Darren J Hsu 외 arxiv

Make-on-demand combinatorial synthesis libraries (CSLs) like Enamine REAL have significantly enabled drug discovery efforts. However, their large size presents a challenge for virtual screening, where the goal is to iden…

Drug Discovery

Efficient Budget Allocation for Large-Scale LLM-Enabled Virtual Screening

2024-08-18 · Zaile Li, Weiwei Fan, L. Jeff Hong

Screening tasks that aim to identify a small subset of top alternatives from a large pool are common in business decision-making processes. These tasks often require substantial human effort to evaluate each alternative'…

Computational Efficiency

SVSBI: Sequence-based virtual screening of biomolecular interactions

2022-12-27 · Li Shen, Hongsong Feng, Yuchi Qiu, Guo-Wei Wei

Virtual screening (VS) is an essential technique for understanding biomolecular interactions, particularly, drug design and discovery. The best-performing VS models depend vitally on three-dimensional (3D) structures, wh…

Drug DesignDrug DiscoveryMolecular Docking