paper-with-me

홈 › Papers

A Dataset for Distilling Knowledge Priors from Literature for Therapeutic Design

2025-08-14 · Haydn Thomas Jones, Natalie Maus, Josh Magnus Ludan, Maggie Ziyu Huan, Jiaming Liang, Marcelo Der Torossian Torres, Jiatao Liang, Zachary Ives, Yoseph Barash, Cesar de la Fuente-Nunez, Jacob R. Gardner, Mark Yatskar arxiv

AI-driven discovery can greatly reduce design time and enhance new therapeutics' effectiveness. Models using simulators explore broad design spaces but risk violating implicit constraints due to a lack of experimental priors. For example, in a new analysis we performed on a diverse set of models on the GuacaMol benchmark using supervised classifiers, over 60\% of molecules proposed had high probability of being mutagenic. In this work, we introduce Medex, a dataset of priors for design problems extracted from literature describing compounds used in lab settings. It is constructed with LLM pipelines for discovering therapeutic entities in relevant paragraphs and summarizing information in concise fair-use facts. Medex consists of 32.3 million pairs of natural language facts, and appropriate entity representations (i.e. SMILES or refseq IDs). To demonstrate the potential of the data, we train LLM, CLIP, and LLava architectures to reason jointly about text and design targets and evaluate on tasks from the Therapeutic Data Commons (TDC). Medex is highly effective for creating models with strong priors: in supervised prediction problems that use our data as pretraining, our best models with 15M learnable parameters outperform larger 2B TxGemma on both regression and classification TDC tasks, and perform comparably to 9B models on average. Models built with Medex can be used as constraints while optimizing for novel molecules in GuacaMol, resulting in proposals that are safer and nearly as effective. We release our dataset at https://huggingface.co/datasets/medexanon/Medex, and will provide expanded versions as available literature grows.

📄 PDF Abstract BibTeX arXiv:2508.10899

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Distilling Textual Priors from LLM to Efficient Image Fusion

2025-04-09 · Ran Zhang, Xuanhua He, Ke Cao, Liu Liu 외

Multi-modality image fusion aims to synthesize a single, comprehensive image from multiple source inputs. Traditional approaches, such as CNNs and GANs, offer efficiency but struggle to handle low-quality or complex inpu…

Computational Efficiency

Distilling Object Detectors with Task Adaptive Regularization

2020-06-23 · Ruoyu Sun, Fuhui Tang, Xiaopeng Zhang, Hongkai Xiong 외

Current state-of-the-art object detectors are at the expense of high computational costs and are hard to deploy to low-end devices. Knowledge distillation, which aims at training a smaller student network by transferring…

Knowledge DistillationObjectRegion Proposal

Tracking Short-Term Temporal Linguistic Dynamics to Characterize Candidate Therapeutics for COVID-19 in the CORD-19 Corpus

2021-01-09 · James Powell, Kari Sentz

Scientific literature tends to grow as a function of funding and interest in a given field. Mining such literature can reveal trends that may not be immediately apparent. The CORD-19 corpus represents a growing corpus of…

Distilling Visual Priors from Self-Supervised Learning

2020-08-01 · Bingchen Zhao, Xin Wen

Convolutional Neural Networks (CNNs) are prone to overfit small training datasets. We present a novel two-phase pipeline that leverages self-supervised learning and knowledge distillation to improve the generalization ab…

ClassificationContrastive LearningGeneral Classificationimage-classification+3

HiddenObjects: Scalable Diffusion-Distilled Spatial Priors for Object Placement

2026-04-12 · Marco Schouten, Ioannis Siglidis, Serge Belongie, Dim P. Papadopoulos arxiv

We propose a method to learn explicit, class-conditioned spatial priors for object placement in natural scenes by distilling the implicit placement knowledge encoded in text-conditioned diffusion models. Prior work relie…

Image Editing