paper-with-me

홈 › Papers

Beyond Manual Curation: Augmenting Targeted Protein Degradation Databases via Agentic Literature Extraction Workflows

2026-05-11 · Yaochen Rao, Farzaneh Jalalypour, N. M. Anoop Krishnan, Rocío Mercado arxiv

Predictive models in biomedicine depend on structured assay data locked in the text, tables, and supplements of primary publications. This bottleneck is especially acute in targeted protein degradation (TPD), where each assay record must combine compound identity, degradation target, recruiter, assay context, and endpoint values reported across sections, tables, and supplementary files. Inconsistent compound identifiers and incomplete or implicit assay context further demand domain-specific logic that generic LLM pipelines do not provide. Existing molecular glue and PROTAC databases are manually curated and often lack the experimental context required for downstream modeling. We formulate TPD database extraction as a domain-specific curation task and present an expert-in-the-loop LLM workflow, evaluated through a triangular comparison among LLM predictions, standardized baseline records, and expert-annotated ground truth. A lightweight cross-validated prompt-refinement module adapts extraction instructions from scarce expert annotations. With only seven annotated molecular glue publications, the workflow achieved record-level $F_1 = 0.98$ and transferred to PROTACs by terminology substitution alone, maintaining record-level $F_1 > 0.93$. Applied at scale, it expanded molecular glue and PROTAC databases by 81% and 92% records, respectively, with 92% and 82.5% of newly recovered records validated as correct upon expert review. The workflow also recovered kinetic and assay-context information essential for cross-study potency comparison and condition-aware degradation modeling. We release the workflow, prompts, evaluation code, and extracted datasets as resources for TPD data curation and AI-assisted scientific curation more broadly.

📄 PDF Abstract BibTeX arXiv:2605.11221

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Best practices for the manual curation of Intrinsically Disordered Proteins in DisProt

2023-10-25 · Federica Quaglia, Anastasia Chasapi, Maria Victoria Nugnes, Maria Cristina Aspromonte 외

The DisProt database is a significant resource containing manually curated data on experimentally validated intrinsically disordered proteins (IDPs) and regions (IDRs) from the literature. Developed in 2005, its primary …

Self Distillation Fine-Tuning of Protein Language Models Improves Versatility in Protein Design

2025-12-10 · Amin Tavakoli, Raswanth Murugan, Ozan Gokdemir, Arvind Ramanathan 외 arxiv

Supervised fine-tuning (SFT) is a standard approach for adapting large language models to specialized domains, yet its application to protein sequence modeling and protein language models (PLMs) remains ad hoc. This is i…

Protein Language ModelProtein Design

Large-scale protein-protein post-translational modification extraction with distant supervision and confidence calibrated BioBERT

2022-01-06 · Aparna Elangovan, Yuan Li, Douglas E. V. Pires, Melissa J. Davis 외

Protein-protein interactions (PPIs) are critical to normal cellular function and are related to many disease pathways. However, only 4% of PPIs are annotated with PTMs in biological knowledge databases such as IntAct, ma…

Translocatome: a novel resource for the analysis of protein translocation between cellular organelles

2018-10-15 · Peter Mendik, Levente Dobronyi, Ferenc Hari, Csaba Kerepesi 외

Here we present Translocatome, the first dedicated database of human translocating proteins. The core of the Translocatome database is the manually curated data set of 213 human translocating proteins listing the source …

An X-ray absorption spectrum database for iron-containing proteins

2025-04-14 · Yufeng Wang, Peiyao Wang, Emerita Mendoza Rengifo, Dali Yang 외

Earth-abundant iron is an essential metal in regulating the structure and function of proteins. This study presents the development of a comprehensive X-ray Absorption Spectroscopy (XAS) database focused on iron-containi…