L+M-24: Building a Dataset for Language + Molecules @ ACL 2024
Language-molecule models have emerged as an exciting direction for molecular discovery and understanding. However, training these models is challenging due to the scarcity of molecule-language pair datasets. At this point, datasets have been released which are 1) small and scraped from existing databases, 2) large but noisy and constructed by performing entity linking on the scientific literature, and 3) built by converting property prediction datasets to natural language using templates. In this document, we detail the $\textit{L+M-24}$ dataset, which has been created for the Language + Molecules Workshop shared task at ACL 2024. In particular, $\textit{L+M-24}$ is designed to focus on three key benefits of natural language in molecule design: compositionality, functionality, and abstraction.
Code (1)
Tasks
Entity LinkingProperty PredictionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
mCLM: A Function-Infused and Synthesis-Friendly Modular Chemical Language Model
Despite their ability to understand chemical knowledge and accurately generate sequential representations, large language models (LLMs) remain limited in their capacity to propose novel molecules with drug-like propertie…
Language ModelingLanguage ModellingSynLlama: Generating Synthesizable Molecules and Their Analogs with Large Language Models
Generative machine learning models for small molecule drug discovery have shown immense promise, but many molecules they generate are too difficult to synthesize, making them impractical for further investigation or deve…
Drug DiscoveryMolecular Graph Generation by Decomposition and Reassembling
Designing molecular structures with desired chemical properties is an essential task in drug discovery and material design. However, finding molecules with the optimized desired properties is still a challenging task due…
Drug DiscoveryGraph GenerationMolecular Graph GenerationvalidKnowMol: Advancing Molecular Large Language Models with Multi-Level Chemical Knowledge
The molecular large language models have garnered widespread attention due to their promising potential on molecular applications. However, current molecular large language models face significant limitations in understa…
Learning Graph Models for Template-Free Retrosynthesis
Retrosynthesis prediction is a fundamental problem in organic synthesis, where the task is to identify precursor molecules that can be used to synthesize a target molecule. A key consideration in building neural models f…
RetrosynthesisSingle-step retrosynthesis