paper-with-me

홈 › Papers

Benchmarking large language models for materials synthesis: the case of atomic layer deposition

2024-12-13 · Angel Yanguas-Gil, Matthew T. Dearing, Jeffrey W. Elam, Jessica C. Jones, Sungjoon Kim, Adnan Mohammad, Chi Thang Nguyen, Bratin Sengupta

In this work we introduce an open-ended question benchmark, ALDbench, to evaluate the performance of large language models (LLMs) in materials synthesis, and in particular in the field of atomic layer deposition, a thin film growth technique used in energy applications and microelectronics. Our benchmark comprises questions with a level of difficulty ranging from graduate level to domain expert current with the state of the art in the field. Human experts reviewed the questions along the criteria of difficulty and specificity, and the model responses along four different criteria: overall quality, specificity, relevance, and accuracy. We ran this benchmark on an instance of OpenAI's GPT-4o. The responses from the model received a composite quality score of 3.7 on a 1 to 5 scale, consistent with a passing grade. However, 36% of the questions received at least one below average score. An in-depth analysis of the responses identified at least five instances of suspected hallucination. Finally, we observed statistically significant correlations between the difficulty of the question and the quality of the response, the difficulty of the question and the relevance of the response, and the specificity of the question and the accuracy of the response as graded by the human experts. This emphasizes the need to evaluate LLMs across multiple criteria beyond difficulty or accuracy.

📄 PDF Abstract BibTeX arXiv:2412.10477

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingHallucinationSpecificity

Similar Papers 제목 키워드 기반

Coupling Language Models with Physics-based Simulation for Synthesis of Inorganic Materials

2026-05-29 · Edward W. Staley, Tom Arbaugh, Michael Pekala, Alexander New 외 arxiv

Modern generative machine learning (ML) models can propose novel inorganic crystalline materials with targeted properties; however, synthesis planning of these materials remains difficult due to the complexity of the ass…

MSQA: Benchmarking LLMs on Graduate-Level Materials Science Reasoning and Knowledge

2025-05-29 · Jerry Junyang Cheung, Shiyao Shen, Yuchen Zhuang, Yinghao Li 외

Despite recent advances in large language models (LLMs) for materials science, there is a lack of benchmarks for evaluating their domain-specific knowledge and complex reasoning abilities. To bridge this gap, we introduc…

Benchmarking

Large Language Model-Guided Prediction Toward Quantum Materials Synthesis

2024-10-28 · Ryotaro Okabe, Zack West, Abhijatmedhi Chotrattanapituk, Mouyang Cheng 외

The synthesis of inorganic crystalline materials is essential for modern technology, especially in quantum materials development. However, designing efficient synthesis workflows remains a significant challenge due to th…

Language ModelingLanguage ModellingLarge Language Model

Agentic Mixture-of-Workflows for Multi-Modal Chemical Search

2025-02-26 · Tiffany J. Callahan, Nathaniel H. Park, Sara Capponi

The vast and complex materials design space demands innovative strategies to integrate multidisciplinary scientific knowledge and optimize materials discovery. While large language models (LLMs) have demonstrated promisi…

BenchmarkingRetrievalRetrieval-augmented Generation

Towards Fully-Automated Materials Discovery via Large-Scale Synthesis Dataset and Expert-Level LLM-as-a-Judge

2025-02-23 · Heegyu Kim, Taeyang Jeon, Seungtaek Choi, Jihoon Hong 외

Materials synthesis is vital for innovations such as energy storage, catalysis, electronics, and biomedical devices. Yet, the process relies heavily on empirical, trial-and-error methods guided by expert intuition. Our w…

Experimental Design