paper-with-me

Papers

DARWIN 1.5: Large Language Models as Materials Science Adapted Learners

2024-12-16 · Tong Xie, Yuwei Wan, Yixuan Liu, Yuchen Zeng, Shaozhou Wang, Wenjie Zhang, Clara Grazian, Chunyu Kit, Wanli Ouyang, Dongzhan Zhou, Bram Hoex

Materials discovery and design aim to find compositions and structures with desirable properties over highly complex and diverse physical spaces. Traditional solutions, such as high-throughput simulations or machine learning, often rely on complex descriptors, which hinder generalizability and transferability across different material systems. Moreover, These descriptors may inadequately represent macro-scale material properties, which are influenced by structural imperfections and compositional variations in real-world samples, thus limiting their practical applicability. To address these challenges, we propose DARWIN 1.5, the largest open-source large language model tailored for materials science. By leveraging natural language as input, DARWIN eliminates the need for task-specific descriptors and enables a flexible, unified approach to material property prediction and discovery. Our approach integrates 6M material domain papers and 21 experimental datasets from 49,256 materials across modalities while enabling cross-task knowledge transfer. The enhanced model achieves up to 59.1% improvement in prediction accuracy over the base LLaMA-7B architecture and outperforms SOTA machine learning approaches across 8 materials design tasks. These results establish LLMs as a promising foundation for developing versatile and scalable models in materials science.

📄 PDF Abstract BibTeX arXiv:2412.11970

Code (1)

masterai-eam/darwin 공식 구현 pytorch

Tasks

Large Language ModelMulti-Task LearningProperty PredictionQuestion AnsweringTransfer Learning

Methods 이 논문이 사용한 방법론

BASE 설명 없음

Similar Papers 제목 키워드 기반

DARWIN Series: Domain Specific Large Language Models for Natural Science

2023-08-25 · Tong Xie, Yuwei Wan, Wei Huang, Zhenyu Yin 외

Emerging tools bring forth fresh approaches to work, and the field of natural science is no different. In natural science, traditional manual, serial, and labour-intensive work is being augmented by automated, parallel, …

Knowledge Graphs

Data Darwinism Part I: Unlocking the Value of Scientific Data for Pre-training

2026-02-08 · Yiwei Qin, Zhen Huang, Tiantian Mi, Weiye Si 외 arxiv

Data quality determines foundation model performance, yet systematic processing frameworks are lacking. We introduce Data Darwinism, a ten-level taxonomy (L0-L9) that conceptualizes data-model co-evolution: advanced mode…

MaterialFigBENCH: benchmark dataset with figures for evaluating college-level materials science problem-solving abilities of multimodal large language models

2026-03-12 · Michiko Yoshitake, Yuta Suzuki, Ryo Igarashi, Yoshitaka Ushiku 외 arxiv

We present MaterialFigBench, a benchmark dataset designed to evaluate the ability of multimodal large language models (LLMs) to solve university-level materials science problems that require accurate interpretation of fi…

Multimodal ReasoningVisual Reasoning

Are LLMs Ready for Real-World Materials Discovery?

2024-02-07 · Santiago Miret, N M Anoop Krishnan

Large Language Models (LLMs) create exciting possibilities for powerful language processing tools to accelerate research in materials science. While LLMs have great potential to accelerate materials understanding and dis…

Interpretable discovery of new semiconductors with machine learning

2021-01-12 · Hitarth Choubisa, Petar Todorović, Joao M. Pina, Darshan H. Parmar 외

Machine learning models of materials$^{1-5}$ accelerate discovery compared to ab initio methods: deep learning models now reproduce density functional theory (DFT)-calculated results at one hundred thousandths of the cos…

BIG-bench Machine LearningKnowledge Distillation