paper-with-me

Papers

TxBench-PP: Analyzing AI Agent Performance on Small-Molecule Preclinical Pharmacology

2026-06-17 · Hannah Le, Ramesh Ramasamy, Alex Urrutia, Mahsa Yazdani, Tim Proctor, Kenny Workman arxiv

Artificial intelligence (AI) agents promise to accelerate drug discovery by compressing interpretation and decision-making loops, but practical deployment requires trusted evaluation on realistic program decisions. We introduce TherapeuticsBench Preclinical Pharmacology (TxBench-PP), a verifiable benchmark for small-molecule preclinical pharmacology and the first focused slice of a broader TherapeuticsBench effort across drug-discovery stages and therapeutic modalities. TxBench-PP tests whether agents can recover accurate conclusions from real-world assay data rather than memorized facts from literature. The benchmark contains 100 evaluations indexed by program stage, assay type, and task structure, spanning mechanism-of-action (MoA) and pharmacodynamic (PD) reasoning, compound-target engagement, causal target validation, developability and safety, and translational efficacy. Agents receive realistic workflow snapshots, inspect files in a coding environment, and return structured answers graded deterministically. Across 16 model-harness configurations, comprising 11 models and 4,800 trajectories, no system reliably recovered preclinical pharmacology decisions. The strongest configuration, Claude Opus 4.8 / Pi, passed 59.3\% of endpoint attempts (178/300; 95\% CI, 51.1-67.6), followed by GPT-5.5 / Pi at 55.3\% (166/300; 47.0-63.6).

📄 PDF Abstract BibTeX arXiv:2606.19245

Code (0)

등록된 구현이 없습니다.

Tasks

Drug Discovery

Similar Papers 제목 키워드 기반

PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX

2026-08-18 · Genghan Zhang, Yixin Dong, Chengze Fan, Zhichen Zeng 외 arxiv

We introduce PTXBench, a benchmark for evaluating and adapting large language models (LLMs) to use architecture-specific PTX for GPU kernel optimization. PTXBench measures functional correctness, whether selected target …

SMDD-Bench: Can LLMs Solve Real-World Small Molecule Drug Design Tasks?

2026-05-20 · Kevin Han, Renfei Zhang, Kathy Wei, Hamed Mahdavi 외 arxiv

LLM agents have incredible potential for scientific discovery applications. However, the performance of LLM agents on real-world, small molecule drug design (SMDD) tasks across diverse chemistries and targets is unclear.…

Question Answering

Linear-scaling kernels for protein sequences and small molecules outperform deep learning while providing uncertainty quantitation and improved interpretability

2023-02-07 · Jonathan Parkinson, Wei Wang

Gaussian process (GP) is a Bayesian model which provides several advantages for regression tasks in machine learning such as reliable quantitation of uncertainty and improved interpretability. Their adoption has been pre…

Data VisualizationDrug DiscoveryFormation EnergyProtein Function Prediction

Molecular dynamics simulations with grand-canonical reweighting suggest cooperativity effects in RNA structure probing experiments

2022-09-26 · Nicola Calonaci, Mattia Bernetti, Alisha Jones, Michael Sattler 외

Chemical probing experiments such as SHAPE are routinely used to probe RNA molecules. In this work, we use atomistic molecular dynamics simulations to test the hypothesis that binding of RNA with SHAPE reagents is affect…

Improving homology-directed repair by small molecule agents for genetic engineering in unconventional yeast? -- Learning from the engineering of mammalian systems

2023-08-29 · Min Lu, Sonja Billerbeck

The ability to precisely edit genomes by deleting or adding genetic information enables the study of biological functions and the building of efficient cell factories. In many unconventional yeasts, such as promising new…