Extracting Structured Seed-Mediated Gold Nanorod Growth Procedures from Literature with GPT-3
Although gold nanorods have been the subject of much research, the pathways for controlling their shape and thereby their optical properties remain largely heuristically understood. Although it is apparent that the simultaneous presence of and interaction between various reagents during synthesis control these properties, computational and experimental approaches for exploring the synthesis space can be either intractable or too time-consuming in practice. This motivates an alternative approach leveraging the wealth of synthesis information already embedded in the body of scientific literature by developing tools to extract relevant structured data in an automated, high-throughput manner. To that end, we present an approach using the powerful GPT-3 language model to extract structured multi-step seed-mediated growth procedures and outcomes for gold nanorods from unstructured scientific text. GPT-3 prompt completions are fine-tuned to predict synthesis templates in the form of JSON documents from unstructured text input with an overall accuracy of $86\%$. The performance is notable, considering the model is performing simultaneous entity recognition and relation extraction. We present a dataset of 11,644 entities extracted from 1,137 papers, resulting in 268 papers with at least one complete seed-mediated gold nanorod growth procedure and outcome for a total of 332 complete procedures.
Code (0)
등록된 구현이 없습니다.
Tasks
Language ModellingRelation ExtractionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Interplay of Electrostatic Interaction and Steric Repulsion between Bacteria and Gold Surface Influences Raman Enhancement
Plasmonic nanostructures have wide applications in photonics including pathogen detection and diagnosis via Surface-Enhanced Raman Spectroscopy (SERS). Despite major role plasmonics play in signal enhancement, electrosta…
A digital microarray using interferometric detection of plasmonic nanorod labels
DNA and protein microarrays are a high-throughput technology that allow the simultaneous quantification of tens of thousands of different biomolecular species. The mediocre sensitivity and dynamic range of traditional fl…
DiagnosticSensitivityExtracting periodontitis diagnosis in clinical notes with RoBERTa and regular expression
This study aimed to utilize text processing and natural language processing (NLP) models to mine clinical notes for the diagnosis of periodontitis and to evaluate the performance of a named entity recognition (NER) model…
named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NERGood Seed Makes a Good Crop: Discovering Secret Seeds in Text-to-Image Diffusion Models
Recent advances in text-to-image (T2I) diffusion models have facilitated creative and photorealistic image synthesis. By varying the random seeds, we can generate many images for a fixed text prompt. Technically, the see…
Image GenerationLeveraging large language models for structured information extraction from pathology reports
Background: Structured information extraction from unstructured histopathology reports facilitates data accessibility for clinical research. Manual extraction by experts is time-consuming and expensive, limiting scalabil…