paper-with-me

홈 › Papers

Automated Annotation of Scientific Texts for ML-based Keyphrase Extraction and Validation

2023-11-08 · Oluwamayowa O. Amusat, Harshad Hegde, Christopher J. Mungall, Anna Giannakou, Neil P. Byers, Dan Gunter, Kjiersten Fagnan, Lavanya Ramakrishnan

Advanced omics technologies and facilities generate a wealth of valuable data daily; however, the data often lacks the essential metadata required for researchers to find and search them effectively. The lack of metadata poses a significant challenge in the utilization of these datasets. Machine learning-based metadata extraction techniques have emerged as a potentially viable approach to automatically annotating scientific datasets with the metadata necessary for enabling effective search. Text labeling, usually performed manually, plays a crucial role in validating machine-extracted metadata. However, manual labeling is time-consuming; thus, there is an need to develop automated text labeling techniques in order to accelerate the process of scientific innovation. This need is particularly urgent in fields such as environmental genomics and microbiome science, which have historically received less attention in terms of metadata curation and creation of gold-standard text mining datasets. In this paper, we present two novel automated text labeling approaches for the validation of ML-generated metadata for unlabeled texts, with specific applications in environmental genomics. Our techniques show the potential of two new ways to leverage existing information about the unlabeled texts and the scientific domain. The first technique exploits relationships between different types of data sources related to the same research study, such as publications and proposals. The second technique takes advantage of domain-specific controlled vocabularies or ontologies. In this paper, we detail applying these approaches for ML-generated metadata validation. Our results show that the proposed label assignment approaches can generate both generic and highly-specific text labels for the unlabeled texts, with up to 44% of the labels matching with those suggested by a ML keyword extraction algorithm.

📄 PDF Abstract BibTeX arXiv:2311.05042

Code (0)

등록된 구현이 없습니다.

Tasks

Keyphrase ExtractionKeyword Extraction

Similar Papers 제목 키워드 기반

Keyphrase Extraction from Scientific Articles via Extractive Summarization

2021-06-01 · NAACL (sdp) 2021 6 · Chrysovalantis Giorgos Kontoulis, Eirini Papagiannopoulou, Grigorios Tsoumakas

Automatically extracting keyphrases from scholarly documents leads to a valuable concise representation that humans can understand and machines can process for tasks, such as information retrieval, article clustering and…

ArticlesExtractive SummarizationInformation RetrievalKeyphrase Extraction+1

Local Word Vectors Guiding Keyphrase Extraction

2017-10-20 · Eirini Papagiannopoulou, Grigorios Tsoumakas

Automated keyphrase extraction is a fundamental textual information processing task concerned with the selection of representative phrases from a document that summarize its content. This work presents a novel unsupervis…

Keyphrase ExtractionWord Embeddings

A Joint Learning Approach based on Self-Distillation for Keyphrase Extraction from Scientific Documents

2020-10-22 · COLING 2020 8 · Tuan Manh Lai, Trung Bui, Doo Soon Kim, Quan Hung Tran

Keyphrase extraction is the task of extracting a small set of phrases that best describe a document. Most existing benchmark datasets for the task typically have limited numbers of annotated documents, making it challeng…

ArticlesKeyphrase Extraction

Exploring Fine-tuned Generative Models for Keyphrase Selection: A Case Study for Russian

2024-09-16 · Anna Glazkova, Dmitry Morozov

Keyphrase selection plays a pivotal role within the domain of scholarly texts, facilitating efficient information retrieval, summarization, and indexing. In this work, we explored how to apply fine-tuned generative trans…

Information RetrievalKeyphrase Extraction

LongKey: Keyphrase Extraction for Long Documents

2024-11-26 · Jeovane Honorio Alves, Radu State, Cinthia Obladen de Almendra Freitas, Jean Paul Barddal

In an era of information overload, manually annotating the vast and growing corpus of documents and scholarly papers is increasingly impractical. Automated keyphrase extraction addresses this challenge by identifying rep…

Keyphrase ExtractionLanguage ModelingLanguage Modelling