paper-with-me

Papers

When Single Answer Is Not Enough: Rethinking Single-Step Retrosynthesis Benchmarks for LLMs

2026-02-03 · Bogdan Zagribelnyy, Ivan Ilin, Maksim Kuznetsov, Nikita Bondarev, Mathieu Reymond, Roman Schutski, Thomas MacDougall, Rim Shayakhmetov, Zulfat Miftakhutdinov, Mikolaj Mizera, Vladimir Aladinskiy, Alex Aliper, Alex Zhavoronkov arxiv

Recent progress has expanded the use of large language models (LLMs) in drug discovery, including synthesis planning. However, objective evaluation of retrosynthesis performance remains limited. Existing benchmarks and metrics typically rely on published synthetic procedures and Top-K accuracy based on single ground-truth, which does not capture the open-ended nature of real-world synthesis planning. We propose a new benchmarking framework for single-step retrosynthesis that evaluates both general-purpose and chemistry-specialized LLMs using ChemCensor, a novel metric for chemical plausibility. By emphasizing plausibility over exact match, this approach better aligns with human synthesis planning practices. We also introduce CREED, a novel dataset comprising millions of ChemCensor-validated reaction records for LLM training, and use it to train a model that improves over the LLM baselines under this benchmark.

📄 PDF Abstract BibTeX arXiv:2602.03554

Code (0)

등록된 구현이 없습니다.

Tasks

Single-step retrosynthesisDrug Discovery

Similar Papers 제목 키워드 기반

Rethinking Evaluation in ASR: Are Our Models Robust Enough?

2020-10-22 · Tatiana Likhomanenko, Qiantong Xu, Vineel Pratap, Paden Tomasello 외

Is pushing numbers on a single benchmark valuable in automatic speech recognition? Research results in acoustic modeling are typically evaluated based on performance on a single dataset. While the research community has …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

DiverseNet: When One Right Answer is not Enough

2020-08-24 · CVPR 2018 6 · Michael Firman, Neill D. F. Campbell, Lourdes Agapito, Gabriel J. Brostow

Many structured prediction tasks in machine vision have a collection of acceptable answers, instead of one definitive ground truth answer. Segmentation of images, for example, is subject to human labeling bias. Similarly…

Multiple-choiceStructured Prediction

Single Character Perturbations Break LLM Alignment

2024-07-03 · Leon Lin, Hannah Brown, Kenji Kawaguchi, Michael Shieh

When LLMs are deployed in sensitive, human-facing settings, it is crucial that they do not output unsafe, biased, or privacy-violating outputs. For this reason, models are both trained and instructed to refuse to answer …

Building and Evaluating Open-Domain Dialogue Corpora with Clarifying Questions

2021-09-13 · EMNLP 2021 11 · Mohammad Aliannejadi, Julia Kiseleva, Aleksandr Chuklin, Jeffrey Dalton 외

Enabling open-domain dialogue systems to ask clarifying questions when appropriate is an important direction for improving the quality of the system response. Namely, for cases when a user request is not specific enough …

One Explanation is Not Enough: Structured Attention Graphs for Image Classification

2020-11-13 · NeurIPS 2021 12 · Vivswan Shitole, Li Fuxin, Minsuk Kahng, Prasad Tadepalli 외

Attention maps are a popular way of explaining the decisions of convolutional networks for image classification. Typically, for each image of interest, a single attention map is produced, which assigns weights to pixels …

ClassificationcounterfactualGeneral Classificationimage-classification+1