paper-with-me

Papers

Explainable Molecular Property Prediction: Aligning Chemical Concepts with Predictions via Language Models

2024-05-25 · Zhenzhong Wang, Zehui Lin, WanYu Lin, Ming Yang, Minggang Zeng, Kay Chen Tan

Providing explainable molecular property predictions is critical for many scientific domains, such as drug discovery and material science. Though transformer-based language models have shown great potential in accurate molecular property prediction, they neither provide chemically meaningful explanations nor faithfully reveal the molecular structure-property relationships. In this work, we develop a framework for explainable molecular property prediction based on language models, dubbed as Lamole, which can provide chemical concepts-aligned explanations. We take a string-based molecular representation -- Group SELFIES -- as input tokens to pretrain and fine-tune our Lamole, as it provides chemically meaningful semantics. By disentangling the information flows of Lamole, we propose combining self-attention weights and gradients for better quantification of each chemically meaningful substructure's impact on the model's output. To make the explanations more faithfully respect the structure-property relationship, we then carefully craft a marginal loss to explicitly optimize the explanations to be able to align with the chemists' annotations. We bridge the manifold hypothesis with the elaborated marginal loss to prove that the loss can align the explanations with the tangent space of the data manifold, leading to concept-aligned explanations. Experimental results over six mutagenicity datasets and one hepatotoxicity dataset demonstrate Lamole can achieve comparable classification accuracy and boost the explanation accuracy by up to 14.3%, being the state-of-the-art in explainable molecular property prediction.

📄 PDF Abstract BibTeX arXiv:2405.16041

Code (0)

등록된 구현이 없습니다.

Tasks

Drug DiscoveryMolecular Property Predictionmolecular representationProperty Prediction

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

ProGReST: Prototypical Graph Regression Soft Trees for Molecular Property Prediction

2022-10-07 · Dawid Rymarczyk, Daniel Dobrowolski, Tomasz Danel

In this work, we propose the novel Prototypical Graph Regression Self-explainable Trees (ProGReST) model, which combines prototype learning, soft decision trees, and Graph Neural Networks. In contrast to other works, our…

Graph RegressionMolecular Property PredictionPredictionProperty Prediction+1

Pure Component Property Estimation Framework Using Explainable Machine Learning Methods

2025-05-14 · Jianfeng Jiao, Xi Gao, Jie Li

Accurate prediction of pure component physiochemical properties is crucial for process integration, multiscale modeling, and optimization. In this work, an enhanced framework for pure component property prediction by usi…

molecular representationProperty Prediction

Enhancing Chemical Explainability Through Counterfactual Masking

2025-08-25 · Łukasz Janisiów, Marek Kochańczyk, Bartosz Zieliński, Tomasz Danel arxiv

Molecular property prediction is a crucial task that guides the design of new compounds, including drugs and materials. While explainable artificial intelligence methods aim to scrutinize model predictions by identifying…

Molecular Property Prediction

Unveiling Molecular Secrets: An LLM-Augmented Linear Model for Explainable and Calibratable Molecular Property Prediction

2024-10-11 · Zhuoran Li, Xu sun, WanYu Lin, Jiannong Cao

Explainable molecular property prediction is essential for various scientific fields, such as drug discovery and material science. Despite delivering intrinsic explainability, linear models struggle with capturing comple…

CPUDimensionality ReductionDrug DiscoveryMolecular Property Prediction+2

MolKD: Distilling Cross-Modal Knowledge in Chemical Reactions for Molecular Property Prediction

2023-05-03 · Liang Zeng, Lanqing Li, Jian Li

How to effectively represent molecules is a long-standing challenge for molecular property prediction and drug discovery. This paper studies this problem and proposes to incorporate chemical domain knowledge, specificall…

Drug DiscoveryMolecular Property PredictionPredictionProperty Prediction