paper-with-me

홈 › Papers

EnzChemRED, a rich enzyme chemistry relation extraction dataset

2024-04-22 · Po-Ting Lai, Elisabeth Coudert, Lucila Aimo, Kristian Axelsen, Lionel Breuza, Edouard de Castro, Marc Feuermann, Anne Morgat, Lucille Pourcel, Ivo Pedruzzi, Sylvain Poux, Nicole Redaschi, Catherine Rivoire, Anastasia Sveshnikova, Chih-Hsuan Wei, Robert Leaman, Ling Luo, Zhiyong Lu, Alan Bridge

Expert curation is essential to capture knowledge of enzyme functions from the scientific literature in FAIR open knowledgebases but cannot keep pace with the rate of new discoveries and new publications. In this work we present EnzChemRED, for Enzyme Chemistry Relation Extraction Dataset, a new training and benchmarking dataset to support the development of Natural Language Processing (NLP) methods such as (large) language models that can assist enzyme curation. EnzChemRED consists of 1,210 expert curated PubMed abstracts in which enzymes and the chemical reactions they catalyze are annotated using identifiers from the UniProt Knowledgebase (UniProtKB) and the ontology of Chemical Entities of Biological Interest (ChEBI). We show that fine-tuning pre-trained language models with EnzChemRED can significantly boost their ability to identify mentions of proteins and chemicals in text (Named Entity Recognition, or NER) and to extract the chemical conversions in which they participate (Relation Extraction, or RE), with average F1 score of 86.30% for NER, 86.66% for RE for chemical conversion pairs, and 83.79% for RE for chemical conversion pairs and linked enzymes. We combine the best performing methods after fine-tuning using EnzChemRED to create an end-to-end pipeline for knowledge extraction from text and apply this to abstracts at PubMed scale to create a draft map of enzyme functions in literature to guide curation efforts in UniProtKB and the reaction knowledgebase Rhea. The EnzChemRED corpus is freely available at https://ftp.expasy.org/databases/rhea/nlp/.

📄 PDF Abstract BibTeX arXiv:2404.14209

Code (0)

등록된 구현이 없습니다.

Tasks

Benchmarkingnamed-entity-recognitionNamed Entity RecognitionNERRelationRelation Extraction

Methods 이 논문이 사용한 방법론

Ontology 설명 없음

Similar Papers 제목 키워드 기반

General Multimodal Protein Design Enables DNA-Encoding of Chemistry

2026-04-06 · Jarrid Rector-Brooks, Théophile Lambert, Marta Skreta, Daniel Roth 외 arxiv

Evolution is an extraordinary engine for enzymatic diversity, yet the chemistry it has explored remains a narrow slice of what DNA can encode. Deep generative models can design new proteins that bind ligands, but none ha…

Protein Design

Improving Enzyme Prediction with Chemical Reaction Equations by Hypergraph-Enhanced Knowledge Graph Embeddings

2026-01-08 · Tengwei Song, Long Yin, Zhen Han, Zhiqiang Xu arxiv

Predicting enzyme-substrate interactions has long been a fundamental problem in biochemistry and metabolic engineering. While existing methods could leverage databases of expert-curated enzyme-substrate pairs for models …

Knowledge Graph Embedding

Multi-Alignment Contrastive Learning for Enzyme--Reaction Retrieval

2025-12-09 · Gengmo Zhou, Feng Yu, Wenda Wang, Zhifeng Gao 외 arxiv

Identifying enzymes that catalyze target biochemical reactions is a key step in computational enzyme discovery and biocatalyst design. Recent representation-learning methods formulate this problem as enzyme--reaction mat…

Contrastive Learning

zERExtractor:An Automated Platform for Enzyme-Catalyzed Reaction Data Extraction from Scientific Literature

2025-07-30 · Rui Zhou, Haohui Ma, Tianle Xin, Lixin Zou 외 arxiv

The rapid expansion of enzyme kinetics literature has outpaced the curation capabilities of major biochemical databases, creating a substantial barrier to AI-driven modeling and knowledge discovery. We present zERExtract…

Relation ExtractionTable RecognitionActive Learning

An Information Extraction and Knowledge Graph Platform for Accelerating Biochemical Discoveries

2019-07-19 · Matteo Manica, Christoph Auer, Valery Weber, Federico Zipoli 외

Information extraction and data mining in biochemical literature is a daunting task that demands resource-intensive computation and appropriate means to scale knowledge ingestion. Being able to leverage this immense sour…