paper-with-me

홈 › Papers

Probing Materials Knowledge in LLMs: From Latent Embeddings to Reliable Predictions

2026-03-02 · Vineeth Venugopal, Soroush Mahjoubi, Elsa Olivetti arxiv

Large language models are increasingly applied to materials science, yet fundamental questions remain about their reliability and knowledge encoding. Evaluating 25 LLMs across four materials science tasks -- over 200 base and fine-tuned configurations -- we find that output modality fundamentally determines model behavior. For symbolic tasks, fine-tuning converges to consistent, verifiable answers with reduced response entropy, while for numerical tasks, fine-tuning improves prediction accuracy but models remain inconsistent across repeated inference runs, limiting their reliability as quantitative predictors. For numerical regression, we find that better performance can be obtained by extracting embeddings directly from intermediate transformer layers than from model text output, revealing an ``LLM head bottleneck,'' though this effect is property- and dataset-dependent. Finally, we present a longitudinal study of GPT model performance in materials science, tracking four models over 18 months and observing 9--43\% performance variation that poses reproducibility challenges for scientific applications.

📄 PDF Abstract BibTeX arXiv:2603.01834

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Sampling Latent Material-Property Information From LLM-Derived Embedding Representations

2024-09-18 · Luke P. J. Gilligan, Matteo Cobelli, Hasan M. Sayeed, Taylor D. Sparks 외

Vector embeddings derived from large language models (LLMs) show promise in capturing latent information from the literature. Interestingly, these can be integrated into material embeddings, potentially useful for data-d…

Efficient Knowledge Probing of Large Language Models by Adapting Pre-trained Embeddings

2025-08-08 · Kartik Sharma, Yiqiao Jin, Rakshit Trivedi, Srijan Kumar arxiv

Large language models (LLMs) acquire knowledge across diverse domains such as science, history, and geography encountered during generative pre-training. However, due to their stochasticity, it is difficult to predict wh…

Harry Potter is Still Here! Probing Knowledge Leakage in Targeted Unlearned Large Language Models via Automated Adversarial Prompting

2025-05-22 · Bang Trinh Tran To, Thai Le

This work presents LURK (Latent UnleaRned Knowledge), a novel framework that probes for hidden retained knowledge in unlearned LLMs through adversarial suffix prompting. LURK automatically generates adversarial prompt su…

Diagnostic

Universal Semantic Embeddings of Chemical Elements for Enhanced Materials Inference and Discovery

2025-02-19 · Yunze Jia, Yuehui Xian, Yangyang Xu, Pengfei Dang 외

We present a framework for generating universal semantic embeddings of chemical elements to advance materials inference and discovery. This framework leverages ElementBERT, a domain-specific BERT-based natural language p…

Bayesian Optimization

Improving Preference Extraction In LLMs By Identifying Latent Knowledge Through Classifying Probes

2025-03-22 · Sharan Maiya, Yinhong Liu, Ramit Debnath, Anna Korhonen

Large Language Models (LLMs) are often used as automated judges to evaluate text, but their effectiveness can be hindered by various unintentional biases. We propose using linear classifying probes, trained by leveraging…

Common Sense Reasoning