paper-with-me

홈 › Papers

MolErr2Fix: Benchmarking LLM Trustworthiness in Chemistry via Modular Error Detection, Localization, Explanation, and Revision

2025-08-26 · Yuyang Wu, Jinhui Ye, Shuhao Zhang, Lu Dai, Yonatan Bisk, Olexandr Isayev arxiv

Large Language Models (LLMs) have shown growing potential in molecular sciences, but they often produce chemically inaccurate descriptions and struggle to recognize or justify potential errors. This raises important concerns about their robustness and reliability in scientific applications. To support more rigorous evaluation of LLMs in chemical reasoning, we present the MolErr2Fix benchmark, designed to assess LLMs on error detection and correction in molecular descriptions. Unlike existing benchmarks focused on molecule-to-text generation or property prediction, MolErr2Fix emphasizes fine-grained chemical understanding. It tasks LLMs with identifying, localizing, explaining, and revising potential structural and semantic errors in molecular descriptions. Specifically, MolErr2Fix consists of 1,193 fine-grained annotated error instances. Each instance contains quadruple annotations, i.e,. (error type, span location, the explanation, and the correction). These tasks are intended to reflect the types of reasoning and verification required in real-world chemical communication. Evaluations of current state-of-the-art LLMs reveal notable performance gaps, underscoring the need for more robust chemical reasoning capabilities. MolErr2Fix provides a focused benchmark for evaluating such capabilities and aims to support progress toward more reliable and chemically informed language models. All annotations and an accompanying evaluation API will be publicly released to facilitate future research.

📄 PDF Abstract BibTeX arXiv:2509.00063

Code (0)

등록된 구현이 없습니다.

Tasks

Text Generation

Similar Papers 제목 키워드 기반

Benchmarking Retrieval-Augmented Generation for Chemistry

2025-05-12 · Xianrui Zhong, Bowen Jin, Siru Ouyang, Yanzhen Shen 외

Retrieval-augmented generation (RAG) has emerged as a powerful framework for enhancing large language models (LLMs) with external knowledge, particularly in scientific domains that demand specialized and dynamic informat…

BenchmarkingRAGRetrievalRetrieval-augmented Generation

Metrics for Benchmarking and Uncertainty Quantification: Quality, Applicability, and a Path to Best Practices for Machine Learning in Chemistry

2020-09-30 · Gaurav Vishwakarma, Aditya Sonpal, Johannes Hachmann

This review aims to draw attention to two issues of concern when we set out to make machine learning work in the chemical and materials domain, i.e., statistical loss function metrics for the validation and benchmarking …

BenchmarkingBIG-bench Machine LearningUncertainty Quantification

Harnessing AtomisticSkills for Agentic Atomistic Research

2026-05-18 · Bowen Deng, Bohan Li, Matthew Cox, Hoje Chun 외 arxiv

Computational materials science and chemistry span vast knowledge domains and fractured software ecosystems. Although large language models (LLMs) have demonstrated research capabilities, scaling monolithic agents to man…

Drug Discovery

MATTERIX: toward a digital twin for robotics-assisted chemistry laboratory automation

2026-01-19 · Kourosh Darvish, Arjun Sohal, Abhijoy Mandal, Hatem Fakhruldeen 외 arxiv

Accelerated materials discovery is critical for addressing global challenges. However, developing new laboratory workflows relies heavily on real-world experimental trials, and this can hinder scalability because of the …

BioNeMo Framework: a modular, high-performance library for AI model development in drug discovery

2024-11-15 · Peter St. John, Dejun Lin, Polina Binder, Malcolm Greaves 외

Artificial Intelligence models encoding biology and chemistry are opening new routes to high-throughput and high-quality in-silico drug development. However, their training increasingly relies on computational scale, wit…

Drug Discovery