paper-with-me

홈 › Papers

Automated Standardization of Legacy Biomedical Metadata Using an Ontology-Constrained LLM Agent

2026-03-10 · Josef Hardi, Martin J. O'Connor, Marcos Martinez-Romero, Jean G. Rosario, Stephen A. Fisher, Mark A. Musen arxiv

Scientific metadata are often incomplete and noncompliant with community standards, limiting dataset findability, interoperability, and reuse. Even when standard metadata reporting guidelines exist, they typically lack machine-actionable representations. Producing FAIR datasets requires encoding metadata standards as machine-actionable templates with rich field specifications and precise value constraints. Recent work has shown that LLMs guided by field names and ontology constraints can improve metadata standardization, but these approaches treat constraints as static text prompts, relying on the model's training knowledge alone. We present an LLM-based metadata standardization system that queries standard reporting guidelines and authoritative biomedical terminology services in real time to retrieve canonically correct standards on demand. We evaluate this approach on 839 legacy metadata records from the Human BioMolecular Atlas Program (HuBMAP) using an expert-curated gold standard for exact-match assessment. Our evaluation shows that augmenting the LLM with real-time tool access consistently improves prediction accuracy over the LLM alone across both ontology-constrained and non-ontology-constrained fields, demonstrating a practical approach to automated standardization of biomedical metadata.

📄 PDF Abstract BibTeX arXiv:2604.08552

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Enhancing Automatic PT Tagging for MEDLINE Citations Using Transformer-Based Models

2025-06-03 · Victor H. Cid, James Mork

We investigated the feasibility of predicting Medical Subject Headings (MeSH) Publication Types (PTs) from MEDLINE citation metadata using pre-trained Transformer-based models BERT and DistilBERT. This study addresses li…

Retrieval

Aligning Biomedical Metadata with Ontologies Using Clustering and Embeddings

2019-03-19 · Rafael S. Gonçalves, Maulik R. Kamdar, Mark A. Musen

The metadata about scientific experiments published in online repositories have been shown to suffer from a high degree of representational heterogeneity---there are often many ways to represent the same type of informat…

Clustering

Enhancing Omics Cohort Discovery for Research on Neurodegeneration through Ontology-Augmented Embedding Models

2025-06-16 · José A. Pardo, Alicia Gómez-Pascual, José T. Palma, Juan A. Botía

The growing volume of omics and clinical data generated for neurodegenerative diseases (NDs) requires new approaches for their curation so they can be ready-to-use in bioinformatics. NeuroEmbed is an approach for the eng…

Question Answering

Using association rule mining and ontologies to generate metadata recommendations from multiple biomedical databases

2019-03-21 · Marcos Martínez-Romero, Martin J. O'Connor, Attila L. Egyedi, Debra Willrett 외

Metadata-the machine-readable descriptions of the data-are increasingly seen as crucial for describing the vast array of biomedical datasets that are currently being deposited in public repositories. While most public re…

New Developments in the Polish Parliamentary Corpus

2020-05-01 · LREC 2020 5 · Maciej Ogrodniczuk, Bart{\l}omiej Nito{\'n}

This short paper presents the current (as of February 2020) state of preparation of the Polish Parliamentary Corpus (PPC){---}an extensive collection of transcripts of Polish parliamentary proceedings dating from 1919 to…