paper-with-me

홈 › Papers

OwkinZero: Accelerating Biological Discovery with AI

2025-08-22 · Nathan Bigaud, Vincent Cabeli, Meltem Gürel, Arthur Pignet, John Klein, Gilles Wainrib, Eric Durand arxiv

While large language models (LLMs) are rapidly advancing scientific research, they continue to struggle with core biological reasoning tasks essential for translational and biomedical discovery. To address this limitation, we created and curated eight comprehensive benchmark datasets comprising over 300,000 verifiable question-and-answer pairs, each targeting critical challenges in drug discovery including target druggability, modality suitability, and drug perturbation effects. Using this resource, we developed the OwkinZero models by post-training open-source LLMs through a Reinforcement Learning from Verifiable Rewards strategy. Our results demonstrate that specialized 8-32B OwkinZero models substantially outperform larger, state-of-the-art commercial LLMs on these biological benchmarks. Remarkably, we uncover evidence of a key aspect of generalization: specialist models trained on a single task consistently outperform their base models on previously unseen tasks. This generalization effect is further amplified in our comprehensive OwkinZero models, which were trained on a mixture of datasets and achieve even broader cross-task improvements. This study represents a significant step toward addressing the biological reasoning blind spot in current LLMs, demonstrating that targeted reinforcement learning on carefully curated data can unlock generalizable performance in specialized models, thereby accelerating AI-driven biological discovery.

📄 PDF Abstract BibTeX arXiv:2508.16315

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningDrug Discovery

Similar Papers 제목 키워드 기반

CellFluxRL: Biologically-Constrained Virtual Cell Modeling via Reinforcement Learning

2026-03-23 · Dongxia Wu, Shiye Su, Yuhui Zhang, Elaine Sui 외 arxiv

Building virtual cells with generative models to simulate cellular behavior in silico is emerging as a promising paradigm for accelerating drug discovery. However, prior image-based generative approaches can produce impl…

Reinforcement LearningDrug Discovery

BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments

2024-05-27 · Yusuf Roohani, Andrew Lee, Qian Huang, Jian Vora 외

Agents based on large language models have shown great potential in accelerating scientific discovery by leveraging their rich background knowledge and reasoning capabilities. In this paper, we introduce BioDiscoveryAgen…

AI AgentBayesian Optimizationscientific discovery

SCORCH2: a generalised heterogeneous consensus model for high-enrichment interaction-based virtual screening

2025-04-05 · bioRxiv 2025 4 · Lin Chen, Vincent Blay, Pedro J. Ballester, Douglas R. Houston

The discovery of effective therapeutics remains a complex, costly, and time-consuming endeavor, characterized by high failure rates and significant resource investments. A central bottleneck in early-stage drug discover…

Drug Discovery

DrugPlayGround: Benchmarking Large Language Models and Embeddings for Drug Discovery

2026-02-11 · Tianyu Liu, Sihan Jiang, Fan Zhang, Kunyang Sun 외 arxiv

Large language models (LLMs) are in the ascendancy for research in drug discovery, offering unprecedented opportunities to reshape drug research by accelerating hypothesis generation, optimizing candidate prioritization,…

Drug Discovery

Evaluating the Potential of Leading Large Language Models in Reasoning Biology Questions

2023-11-05 · Xinyu Gong, Jason Holmes, Yiwei Li, Zhengliang Liu 외

Recent advances in Large Language Models (LLMs) have presented new opportunities for integrating Artificial General Intelligence (AGI) into biological research and education. This study evaluated the capabilities of lead…

Logical ReasoningMultiple-choice