paper-with-me

홈 › Papers

Use of a Structured Knowledge Base Enhances Metadata Curation by Large Language Models

2024-04-08 · Sowmya S. Sundaram, Benjamin Solomon, Avani Khatri, Anisha Laumas, Purvesh Khatri, Mark A. Musen

Metadata play a crucial role in ensuring the findability, accessibility, interoperability, and reusability of datasets. This paper investigates the potential of large language models (LLMs), specifically GPT-4, to improve adherence to metadata standards. We conducted experiments on 200 random data records describing human samples relating to lung cancer from the NCBI BioSample repository, evaluating GPT-4's ability to suggest edits for adherence to metadata standards. We computed the adherence accuracy of field name-field value pairs through a peer review process, and we observed a marginal average improvement in adherence to the standard data dictionary from 79% to 80% (p<0.5). We then prompted GPT-4 with domain information in the form of the textual descriptions of CEDAR templates and recorded a significant improvement to 97% from 79% (p<0.01). These results indicate that, while LLMs may not be able to correct legacy metadata to ensure satisfactory adherence to standards when unaided, they do show promise for use in automated metadata curation when integrated with a structured knowledge base

📄 PDF Abstract BibTeX arXiv:2404.05893

Code (1)

musen-lab/biosamplegptcorrection 공식 구현

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Residual Connection 설명 없음
Multi-Head Attention 설명 없음
Adam 설명 없음

Similar Papers 제목 키워드 기반

A Prompt-Aware Structuring Framework for Reliable Reuse of AI-Generated Content in the Agentic Web

2026-05-10 · Shusaku Egami, Masahiro Hamasaki arxiv

The evolution of Large Language Models (LLMs) and the software agents built on them (AI agents) marks a turning point in the transition from a human-centric Web to an ``Agentic Web'' driven by AI agents. However, for AI-…

Knowledge Distillation

Utilising a Large Language Model to Annotate Subject Metadata: A Case Study in an Australian National Research Data Catalogue

2023-10-17 · Shiwei Zhang, Mingfang Wu, Xiuzhen Zhang

In support of open and reproducible research, there has been a rapidly increasing number of datasets made available for research. As the availability of datasets increases, it becomes more important to have quality metad…

In-Context LearningLanguage ModelingLanguage ModellingLarge Language Model

ReLeVAnT: Relevance Lexical Vectors for Accurate Legal Text Classification

2026-04-24 · Ishaan Gakhar, Harsh Nandwani arxiv

The classification of legal documents from an unstructured data corpus has several crucial applications in downstream tasks. Documents relevant to court filings are key in use cases such as drafting motions, memos, and o…

Binary ClassificationText ClassificationKeyword Extraction

User Manual of Automatic Data Curation Tool(ADCT): A bulk data curator software in Library and Information Science

2022-10-31 · A. Banerjee, B. Sutradhar

In library and information science, document storage and user-specific document retrieval are the main aspects of digital library services. To preserve the cultural heritage, documents, and literature, we need a common p…

DescriptiveRetrieval

Automated Metadata Harmonization Using Entity Resolution & Contextual Embedding

2020-10-17 · Kunal Sawarkar, Meenkakshi Kodati

ML Data Curation process typically consist of heterogeneous & federated source systems with varied schema structures; requiring curation process to standardize metadata from different schemas to an inter-operable schema.…

Entity Resolution