paper-with-me

Papers

A Library Perspective on Nearly-Unsupervised Information Extraction Workflows in Digital Libraries

2022-05-02 · Hermann Kroll, Jan Pirklbauer, Florian Plötzky, Wolf-Tilo Balke

Information extraction can support novel and effective access paths for digital libraries. Nevertheless, designing reliable extraction workflows can be cost-intensive in practice. On the one hand, suitable extraction methods rely on domain-specific training data. On the other hand, unsupervised and open extraction methods usually produce not-canonicalized extraction results. This paper tackles the question how digital libraries can handle such extractions and if their quality is sufficient in practice. We focus on unsupervised extraction workflows by analyzing them in case studies in the domains of encyclopedias (Wikipedia), pharmacy and political sciences. We report on opportunities and limitations. Finally we discuss best practices for unsupervised extraction workflows.

📄 PDF Abstract BibTeX arXiv:2205.00716

Code (1)

hermannkroll/kgextractiontoolbox 공식 구현

Similar Papers 제목 키워드 기반

Enhancing Unsupervised Keyword Extraction in Academic Papers through Integrating Highlights with Abstract

2026-04-21 · Yi Xiang, Chengzhi Zhang arxiv

Automatic keyword extraction from academic papers is a key area of interest in natural language processing and information retrieval. Although previous research has mainly focused on utilizing abstract and references for…

Information RetrievalKeyword Extraction

Nearly-Unsupervised Hashcode Representations for Relation Extraction

2019-09-09 · Sahil Garg, Aram Galstyan, Greg Ver Steeg, Guillermo Cecchi

Recently, kernelized locality sensitive hashcodes have been successfully employed as representations of natural language text, especially showing high relevance to biomedical relation extraction tasks. In this paper, we …

RelationRelation Extraction

Nearly-Unsupervised Hashcode Representations for Biomedical Relation Extraction

2019-11-01 · IJCNLP 2019 11 · Sahil Garg, Aram Galstyan, Greg Ver Steeg, Guillermo Cecchi

Recently, kernelized locality sensitive hashcodes have been successfully employed as representations of natural language text, especially showing high relevance to biomedical relation extraction tasks. In this paper, we …

RelationRelation Extraction

An Improved Method for Class-specific Keyword Extraction: A Case Study in the German Business Registry

2024-07-19 · Stephen Meisenbacher, Tim Schopf, Weixin Yan, Patrick Holl 외

The task of $\textit{keyword extraction}$ is often an important initial step in unsupervised information extraction, forming the basis for tasks such as topic modeling or document classification. While recent methods hav…

Document ClassificationKeyword Extraction

Document Intelligence Metrics for Visually Rich Document Evaluation

2022-05-23 · Jonathan Degange, Swapnil Gupta, Zhuoyu Han, Krzysztof Wilkosz 외

The processing of Visually-Rich Documents (VRDs) is highly important in information extraction tasks associated with Document Intelligence. We introduce DI-Metrics, a Python library devoted to VRD model evaluation compri…

Document AI