paper-with-me

홈 › Papers

MaterioMiner -- An ontology-based text mining dataset for extraction of process-structure-property entities

2024-08-05 · Ali Riza Durmaz, Akhil Thomas, Lokesh Mishra, Rachana Niranjan Murthy, Thomas Straub

While large language models learn sound statistical representations of the language and information therein, ontologies are symbolic knowledge representations that can complement the former ideally. Research at this critical intersection relies on datasets that intertwine ontologies and text corpora to enable training and comprehensive benchmarking of neurosymbolic models. We present the MaterioMiner dataset and the linked materials mechanics ontology where ontological concepts from the mechanics of materials domain are associated with textual entities within the literature corpus. Another distinctive feature of the dataset is its eminently fine-granular annotation. Specifically, 179 distinct classes are manually annotated by three raters within four publications, amounting to a total of 2191 entities that were annotated and curated. Conceptual work is presented for the symbolic representation of causal composition-process-microstructure-property relationships. We explore the annotation consistency between the three raters and perform fine-tuning of pre-trained models to showcase the feasibility of named-entity recognition model training. Reusing the dataset can foster training and benchmarking of materials language models, automated ontology construction, and knowledge graph generation from textual data.

📄 PDF Abstract BibTeX arXiv:2408.04661

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingGraph Generationnamed-entity-recognitionNamed Entity Recognition

Methods 이 논문이 사용한 방법론

Ontology 설명 없음

Similar Papers 제목 키워드 기반

Ontology Enrichment by Extracting Hidden Assertional Knowledge from Text

2013-08-03 · Meisam Booshehri, Abbas Malekpour, Peter Luksch, Kamran Zamanifar 외

In this position paper we present a new approach for discovering some special classes of assertional knowledge in the text by using large RDF repositories, resulting in the extraction of new non-taxonomic ontological rel…

PositionRelation Extraction

A Heuristically Modified FP-Tree for Ontology Learning with Applications in Education

2019-10-29 · Safwan Shatnawi, Mohamed Medhat Gaber, Mihaela Cocea

We propose a heuristically modified FP-Tree for ontology learning from text. Unlike previous research, for concept extraction, we use a regular expression parser approach widely adopted in compiler construction, i.e., de…

Question Answeringtext similarity

LLM-based Zero-shot Triple Extraction for Automated Ontology Generation from Software Engineering Standards

2025-08-29 · Songhui Yue arxiv

Ontologies have supported knowledge representation and white-box reasoning for decades; thus, the automated ontology generation (AOG) plays a crucial role in scaling their use. Software engineering standards (SES) consis…

Term Extraction

EPPC-OASIS: Ontology-Aware Adaptation and Structured Inference Refinement for Electronic Patient-Provider Communication Mining in Secure Messages

2026-05-22 · Samah Fodeh, Sreeraj Ramachandran, Elyas Irankhah, Muhammad Arif 외 arxiv

Secure patient-provider messages contain clinically important communication behaviors that are difficult to characterize manually at scale. The Electronic Patient-Provider Communication (EPPC) framework provides an ontol…

Dependency parsing for interaction detection in pharmacogenomics

2012-05-01 · LREC 2012 5 · Gerold Schneider, Fabio Rinaldi, Simon Clematide

We give an overview of our approach to the extraction of interactions between pharmacogenomic entities like drugs, genes and diseases and suggest classes of interaction types driven by data from PharmGKB and partly follo…

Dependency Parsing