paper-with-me

Papers

A Benchmarking Dataset with 2440 Organic Molecules for Volume Distribution at Steady State

2022-11-10 · Wenwen Liu, Cheng Luo, Hecheng Wang, Fanwang Meng

Background: The volume of distribution at steady state (VDss) is a fundamental pharmacokinetics (PK) property of drugs, which measures how effectively a drug molecule is distributed throughout the body. Along with the clearance (CL), it determines the half-life and, therefore, the drug dosing interval. However, the molecular data size limits the generalizability of the reported machine learning models. Objective: This study aims to provide a clean and comprehensive dataset for human VDss as the benchmarking data source, fostering and benefiting future predictive studies. Moreover, several predictive models were also built with machine learning regression algorithms. Methods: The dataset was curated from 13 publicly accessible data sources and the DrugBank database entirely from intravenous drug administration and then underwent extensive data cleaning. The molecular descriptors were calculated with Mordred, and feature selection was conducted for constructing predictive models. Five machine learning methods were used to build regression models, grid search was used to optimize hyperparameters, and ten-fold cross-validation was used to evaluate the model. Results: An enriched dataset of VDss (https://github.com/da-wen-er/VDss) was constructed with 2440 molecules. Among the prediction models, the LightGBM model was the most stable and had the best internal prediction ability with Q2 = 0.837, R2=0.814 and for the other four models, Q2 was higher than 0.79. Conclusions: To the best of our knowledge, this is the largest dataset for VDss, which can be used as the benchmark for computational studies of VDss. Moreover, the regression models reported within this study can be of use for pharmacokinetic related studies.

📄 PDF Abstract BibTeX arXiv:2211.05661

Code (1)

da-wen-er/vdss 공식 구현

Tasks

Benchmarkingfeature selectionregression

Methods 이 논문이 사용한 방법론

Feature Selection Feature selection, also known as variable selection, attribute selection or variable subset selection, is the process of selecting a subset of relevant features (variables,…

Similar Papers 제목 키워드 기반

Alchemy: A Quantum Chemistry Dataset for Benchmarking AI Models

2019-06-22 · Guangyong Chen, Pengfei Chen, Chang-Yu Hsieh, Chee-Kong Lee 외

We introduce a new molecular dataset, named Alchemy, for developing machine learning models useful in chemistry and material science. As of June 20th 2019, the dataset comprises of 12 quantum mechanical properties of 119…

BenchmarkingBIG-bench Machine LearningDiversityGraph Neural Network

CHILI: Chemically-Informed Large-scale Inorganic Nanomaterials Dataset for Advancing Graph Machine Learning

2024-02-20 · Ulrik Friis-Jensen, Frederik L. Johansen, Andy S. Anker, Erik B. Dam 외

Advances in graph machine learning (ML) have been driven by applications in chemistry as graphs have remained the most expressive representations of molecules. While early graph ML methods focused primarily on small orga…

Atomic number classificationBenchmarkingCrystal system classificationDistance regression+9

QuantumChem-200K: A Large-Scale Open Organic Molecular Dataset for Quantum-Chemistry Property Screening and Language Model Benchmarking

2025-11-23 · Yinqi Zeng, Renjie Li arxiv

The discovery of next-generation photoinitiators for two-photon polymerization (TPP) is hindered by the absence of large, open datasets containing the quantum-chemical and photophysical properties required to model photo…

Fast Organic Crystal Structure Prediction with Unit Cell Flow Matching

2026-06-02 · Alston Lo, Luka Mucko, Austin H. Cheng, Andy Cai 외 arxiv

Organic crystal structure prediction (CSP) is a requirement for computational modelling of organic solids, but traditionally costs several CPU-years per molecule. Generative models such as OXtal dramatically reduce this …

An Extendible, Graph-Neural-Network-Based Approach for Accurate Force Field Development of Large Flexible Organic Molecules

2021-06-02 · Xufei Wang, Yuanda Xu, Han Zheng, Kuang Yu

An accurate force field is the key to the success of all molecular mechanics simulations on organic polymers and biomolecules. Accuracy beyond density functional theory is often needed to describe the intermolecular inte…

Graph Neural Network