paper-with-me

홈 › Papers

A general-purpose material property data extraction pipeline from large polymer corpora using Natural Language Processing

2022-09-27 · Pranav Shetty, Arunkumar Chitteth Rajan, Christopher Kuenneth, Sonkakshi Gupta, Lakshmi Prerana Panchumarti, Lauren Holm, Chao Zhang, Rampi Ramprasad

The ever-increasing number of materials science articles makes it hard to infer chemistry-structure-property relations from published literature. We used natural language processing (NLP) methods to automatically extract material property data from the abstracts of polymer literature. As a component of our pipeline, we trained MaterialsBERT, a language model, using 2.4 million materials science abstracts, which outperforms other baseline models in three out of five named entity recognition datasets when used as the encoder for text. Using this pipeline, we obtained ~300,000 material property records from ~130,000 abstracts in 60 hours. The extracted data was analyzed for a diverse range of applications such as fuel cells, supercapacitors, and polymer solar cells to recover non-trivial insights. The data extracted through our pipeline is made available through a web platform at https://polymerscholar.org which can be used to locate material property data recorded in abstracts conveniently. This work demonstrates the feasibility of an automatic pipeline that starts from published literature and ends with a complete set of extracted material property information.

📄 PDF Abstract BibTeX arXiv:2209.13136

Code (1)

Ramprasad-Group/polymer_information_extraction 공식 구현 pytorch

Tasks

ArticlesLanguage ModelingLanguage Modellingnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)

Similar Papers 제목 키워드 기반

LLM4Mat-Bench: Benchmarking Large Language Models for Materials Property Prediction

2024-10-31 · Andre Niyongabo Rubungo, Kangming Li, Jason Hattrick-Simpers, Adji Bousso Dieng

Large language models (LLMs) are increasingly being used in materials science. However, little attention has been given to benchmarking and standardized evaluation for LLM-based materials property prediction, which hinde…

BenchmarkingPredictionProperty Prediction

Flexible, Model-Agnostic Method for Materials Data Extraction from Text Using General Purpose Language Models

2023-02-09 · Maciej P. Polak, Shrey Modi, Anna Latosinska, Jinming Zhang 외

Accurate and comprehensive material databases extracted from research papers are crucial for materials science and engineering, but their development requires significant human effort. With large language models (LLMs) t…

Dynamic In-context Learning with Conversational Models for Data Extraction and Materials Property Prediction

2024-05-16 · Chinedu Ekuma

The advent of natural language processing and large language models (LLMs) has revolutionized the extraction of data from unstructured scholarly papers. However, ensuring data trustworthiness remains a significant challe…

In-Context LearningProperty Prediction

A Survey of AI for Materials Science: Foundation Models, LLM Agents, Datasets, and Tools

2025-06-25 · Minh-Hao Van, Prateek Verma, Chen Zhao, Xintao Wu

Foundation models (FMs) are catalyzing a transformative shift in materials science (MatSci) by enabling scalable, general-purpose, and multimodal AI systems for scientific discovery. Unlike traditional machine learning m…

Continual LearningDomain GeneralizationLarge Language ModelProperty Prediction+1

Large language model-enabled automated data extraction for concrete materials informatics

2026-04-24 · Zhanzhao Li, Kengran Yang, Qiyao He, Kai Gong arxiv

The promise of data-driven materials discovery remains constrained by the scarcity of large, high-quality, and accessible experimental datasets. Here, we introduce a generalizable large language model (LLM)-powered pipel…