Use of a Structured Knowledge Base Enhances Metadata Curation by Large Language Models
Metadata play a crucial role in ensuring the findability, accessibility, interoperability, and reusability of datasets. This paper investigates the potential of large language models (LLMs), specifically GPT-4, to improve adherence to metadata standards. We conducted experiments on 200 random data records describing human samples relating to lung cancer from the NCBI BioSample repository, evaluating GPT-4's ability to suggest edits for adherence to metadata standards. We computed the adherence accuracy of field name-field value pairs through a peer review process, and we observed a marginal average improvement in adherence to the standard data dictionary from 79% to 80% (p<0.5). We then prompted GPT-4 with domain information in the form of the textual descriptions of CEDAR templates and recorded a significant improvement to 97% from 79% (p<0.01). These results indicate that, while LLMs may not be able to correct legacy metadata to ensure satisfactory adherence to standards when unaided, they do show promise for use in automated metadata curation when integrated with a structured knowledge base
Code (1)
Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
A Prompt-Aware Structuring Framework for Reliable Reuse of AI-Generated Content in the Agentic Web
The evolution of Large Language Models (LLMs) and the software agents built on them (AI agents) marks a turning point in the transition from a human-centric Web to an ``Agentic Web'' driven by AI agents. However, for AI-…
Knowledge DistillationUtilising a Large Language Model to Annotate Subject Metadata: A Case Study in an Australian National Research Data Catalogue
In support of open and reproducible research, there has been a rapidly increasing number of datasets made available for research. As the availability of datasets increases, it becomes more important to have quality metad…
In-Context LearningLanguage ModelingLanguage ModellingLarge Language ModelReLeVAnT: Relevance Lexical Vectors for Accurate Legal Text Classification
The classification of legal documents from an unstructured data corpus has several crucial applications in downstream tasks. Documents relevant to court filings are key in use cases such as drafting motions, memos, and o…
Binary ClassificationText ClassificationKeyword ExtractionUser Manual of Automatic Data Curation Tool(ADCT): A bulk data curator software in Library and Information Science
In library and information science, document storage and user-specific document retrieval are the main aspects of digital library services. To preserve the cultural heritage, documents, and literature, we need a common p…
DescriptiveRetrievalAutomated Metadata Harmonization Using Entity Resolution & Contextual Embedding
ML Data Curation process typically consist of heterogeneous & federated source systems with varied schema structures; requiring curation process to standardize metadata from different schemas to an inter-operable schema.…
Entity Resolution