paper-with-me

홈 › Papers

Evaluating the Performance and Robustness of LLMs in Materials Science Q&A and Property Predictions

2024-09-22 · Hongchen Wang, Kangming Li, Scott Ramsay, Yao Fehlis, Edward Kim, Jason Hattrick-Simpers

Large Language Models (LLMs) have the potential to revolutionize scientific research, yet their robustness and reliability in domain-specific applications remain insufficiently explored. In this study, we evaluate the performance and robustness of LLMs for materials science, focusing on domain-specific question answering and materials property prediction across diverse real-world and adversarial conditions. Three distinct datasets are used in this study: 1) a set of multiple-choice questions from undergraduate-level materials science courses, 2) a dataset including various steel compositions and yield strengths, and 3) a band gap dataset, containing textual descriptions of material crystal structures and band gap values. The performance of LLMs is assessed using various prompting strategies, including zero-shot chain-of-thought, expert prompting, and few-shot in-context learning. The robustness of these models is tested against various forms of 'noise', ranging from realistic disturbances to intentionally adversarial manipulations, to evaluate their resilience and reliability under real-world conditions. Additionally, the study showcases unique phenomena of LLMs during predictive tasks, such as mode collapse behavior when the proximity of prompt examples is altered and performance recovery from train/test mismatch. The findings aim to provide informed skepticism for the broad use of LLMs in materials science and to inspire advancements that enhance their robustness and reliability for practical applications.

📄 PDF Abstract BibTeX arXiv:2409.14572

Code (0)

등록된 구현이 없습니다.

Tasks

Band GapIn-Context LearningMultiple-choiceProperty PredictionQuestion Answering

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

MSQA: Benchmarking LLMs on Graduate-Level Materials Science Reasoning and Knowledge

2025-05-29 · Jerry Junyang Cheung, Shiyao Shen, Yuchen Zhuang, Yinghao Li 외

Despite recent advances in large language models (LLMs) for materials science, there is a lack of benchmarks for evaluating their domain-specific knowledge and complex reasoning abilities. To bridge this gap, we introduc…

Benchmarking

LLM4Mat-Bench: Benchmarking Large Language Models for Materials Property Prediction

2024-10-31 · Andre Niyongabo Rubungo, Kangming Li, Jason Hattrick-Simpers, Adji Bousso Dieng

Large language models (LLMs) are increasingly being used in materials science. However, little attention has been given to benchmarking and standardized evaluation for LLM-based materials property prediction, which hinde…

BenchmarkingPredictionProperty Prediction

HalluMat: Detecting Hallucinations in LLM-Generated Materials Science Content Through Multi-Stage Verification

2025-12-26 · Bhanu Prakash Vangala, Sajid Mahmud, Pawan Neupane, Joel Selvaraj 외 arxiv

Artificial Intelligence (AI), particularly Large Language Models (LLMs), is transforming scientific discovery, enabling rapid knowledge generation and hypothesis formulation. However, a critical challenge is hallucinatio…

Probing Materials Knowledge in LLMs: From Latent Embeddings to Reliable Predictions

2026-03-02 · Vineeth Venugopal, Soroush Mahjoubi, Elsa Olivetti arxiv

Large language models are increasingly applied to materials science, yet fundamental questions remain about their reliability and knowledge encoding. Evaluating 25 LLMs across four materials science tasks -- over 200 bas…

MaterialFigBENCH: benchmark dataset with figures for evaluating college-level materials science problem-solving abilities of multimodal large language models

2026-03-12 · Michiko Yoshitake, Yuta Suzuki, Ryo Igarashi, Yoshitaka Ushiku 외 arxiv

We present MaterialFigBench, a benchmark dataset designed to evaluate the ability of multimodal large language models (LLMs) to solve university-level materials science problems that require accurate interpretation of fi…

Multimodal ReasoningVisual Reasoning