paper-with-me

홈 › Papers

ChemPro: A Progressive Chemistry Benchmark for Large Language Models

2026-02-03 · Aaditya Baranwal, Shruti Vyas arxiv

We introduce ChemPro, a progressive benchmark with 4100 natural language question-answer pairs in Chemistry, across 4 coherent sections of difficulty designed to assess the proficiency of Large Language Models (LLMs) in a broad spectrum of general chemistry topics. We include Multiple Choice Questions and Numerical Questions spread across fine-grained information recall, long-horizon reasoning, multi-concept questions, problem-solving with nuanced articulation, and straightforward questions in a balanced ratio, effectively covering Bio-Chemistry, Inorganic-Chemistry, Organic-Chemistry and Physical-Chemistry. ChemPro is carefully designed analogous to a student's academic evaluation for basic to high-school chemistry. A gradual increase in the question difficulty rigorously tests the ability of LLMs to progress from solving basic problems to solving more sophisticated challenges. We evaluate 45+7 state-of-the-art LLMs, spanning both open-source and proprietary variants, and our analysis reveals that while LLMs perform well on basic chemistry questions, their accuracy declines with different types and levels of complexity. These findings highlight the critical limitations of LLMs in general scientific reasoning and understanding and point towards understudied dimensions of difficulty, emphasizing the need for more robust methodologies to improve LLMs.

📄 PDF Abstract BibTeX arXiv:2602.03108

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

PRESTO: Progressive Pretraining Enhances Synthetic Chemistry Outcomes

2024-06-19 · He Cao, Yanjun Shao, Zhiyuan Liu, Zijing Liu 외

Multimodal Large Language Models (MLLMs) have seen growing adoption across various scientific disciplines. These advancements encourage the investigation of molecule-text modeling within synthetic chemistry, a field dedi…

cross-modal alignment

End-to-End Models for Chemical-Protein Interaction Extraction: Better Tokenization and Span-Based Pipeline Strategies

2023-04-03 · Xuguang Ai, Ramakanth Kavuluru

End-to-end relation extraction (E2ERE) is an important task in information extraction, more so for biomedicine as scientific literature continues to grow exponentially. E2ERE typically involves identifying entities (or n…

Chemical-Protein Interaction Extractionnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+4

ChatGPT Chemistry Assistant for Text Mining and Prediction of MOF Synthesis

2023-06-20 · Zhiling Zheng, Oufan Zhang, Christian Borgs, Jennifer T. Chayes 외

We use prompt engineering to guide ChatGPT in the automation of text mining of metal-organic frameworks (MOFs) synthesis conditions from diverse formats and styles of the scientific literature. This effectively mitigates…

ArticlesChatbotPrompt Engineering

ANCHOR-RE: An Agentic Neuro-Symbolic Framework for Grounded Biomedical Relation Extraction

2026-08-04 · Shufan Ming, Yikun Han, Gibong Hong, Rui Zhang 외 arxiv

Biomedical relation extraction (BioRE) extracts structured knowledge from biomedical literature for applications such as knowledge base construction and hypothesis generation. Traditional symbolic systems such as SemRep …

Relation Extraction

ElemeNet: Multiscale Molecular Machine Learning with Uncertainty Quantification Across the Periodic Table

2026-06-29 · Jacob W. Toney, Samir Darouich, Yiran Wang, Aaron G. Garrison 외 arxiv

Advances in deep learning architectures and representations have enabled ML-driven chemical property prediction, but state-of-the-art (SOTA) models have remained largely confined to independent codebases and lack support…