paper-with-me

Papers

Multimodal Approach for Metadata Extraction from German Scientific Publications

2021-11-10 · Azeddine Bouabdallah, Jorge Gavilan, Jennifer Gerbl, Prayuth Patumcharoenpol

Nowadays, metadata information is often given by the authors themselves upon submission. However, a significant part of already existing research papers have missing or incomplete metadata information. German scientific papers come in a large variety of layouts which makes the extraction of metadata a non-trivial task that requires a precise way to classify the metadata extracted from the documents. In this paper, we propose a multimodal deep learning approach for metadata extraction from scientific papers in the German language. We consider multiple types of input data by combining natural language processing and image vision processing. This model aims to increase the overall accuracy of metadata extraction compared to other state-of-the-art approaches. It enables the utilization of both spatial and contextual features in order to achieve a more reliable extraction. Our model for this approach was trained on a dataset consisting of around 8800 documents and is able to obtain an overall F1-score of 0.923.

📄 PDF Abstract BibTeX arXiv:2111.05736

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal Deep Learning

Similar Papers 제목 키워드 기반

MexPub: Deep Transfer Learning for Metadata Extraction from German Publications

2021-06-04 · Zeyd Boukhers, Nada Beili, Timo Hartmann, Prantik Goswami 외

Extracting metadata from scientific papers can be considered a solved problem in NLP due to the high accuracy of state-of-the-art methods. However, this does not apply to German scientific publications, which have a vari…

Transfer Learning

Comparison of Feature Learning Methods for Metadata Extraction from PDF Scholarly Documents

2025-01-09 · Zeyd Boukhers, Cong Yang

The availability of metadata for scientific documents is pivotal in propelling scientific knowledge forward and for adhering to the FAIR principles (i.e. Findability, Accessibility, Interoperability, and Reusability) of …

SMAuC -- The Scientific Multi-Authorship Corpus

2022-11-04 · Janek Bevendorff, Philipp Sauer, Lukas Gienapp, Wolfgang Kircheis 외

The rapidly growing volume of scientific publications offers an interesting challenge for research on methods for analyzing the authorship of documents with one or more authors. However, most existing datasets lack scien…

Automated Annotation of Scientific Texts for ML-based Keyphrase Extraction and Validation

2023-11-08 · Oluwamayowa O. Amusat, Harshad Hegde, Christopher J. Mungall, Anna Giannakou 외

Advanced omics technologies and facilities generate a wealth of valuable data daily; however, the data often lacks the essential metadata required for researchers to find and search them effectively. The lack of metadata…

Keyphrase ExtractionKeyword Extraction

STEREO: Scientific Text Reuse in Open Access Publications

2021-12-22 · Lukas Gienapp, Wolfgang Kircheis, Bjarne Sievers, Benno Stein 외

We present the Webis-STEREO-21 dataset, a massive collection of Scientific Text Reuse in Open-access publications. It contains more than 91 million cases of reused text passages found in 4.2 million unique open-access pu…