paper-with-me

Papers

Using Supervised Learning to Classify Metadata of Research Data by Discipline of Research

2019-10-16 · Tobias Weber, Dieter Kranzlmüller, Michael Fromm, Nelson Tavares de Sousa

Automated classification of metadata of research data by their discipline(s) of research can be used in scientometric research, by repository service providers, and in the context of research data aggregation services. Openly available metadata of the DataCite index for research data were used to compile a large training and evaluation set comprised of 609,524 records, which is published alongside this paper. These data allow to reproducibly assess classification approaches, such as tree-based models and neural networks. According to our experiments with 20 base classes (multi-label classification), multi-layer perceptron models perform best with a f1-macro score of 0.760 closely followed by Long Short-Term Memory models (f1-macro score of 0.755). A possible application of the trained classification models is the quantitative analysis of trends towards interdisciplinarity of digital scholarly output or the characterization of growth patterns of research data, stratified by discipline of research. Both applications perform at scale with the proposed models which are available for re-use.

📄 PDF Abstract BibTeX arXiv:1910.09313

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationGeneral ClassificationMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATION

Similar Papers 제목 키워드 기반

MotifClass: Weakly Supervised Text Classification with Higher-order Metadata Information

2021-11-07 · Yu Zhang, Shweta Garg, Yu Meng, Xiusi Chen 외

We study the problem of weakly supervised text classification, which aims to classify text documents into a set of pre-defined categories with category surface names only and without any annotated training document provi…

text-classificationText Classification

Metadata Enrichment of Multi-Disciplinary Digital Library: A Semantic-based Approach

2018-06-21 · Hussein T. Al-Natsheh, Lucie Martinet, Fabrice Muhlenbach, Fabien Rico 외

In the scientific digital libraries, some papers from different research communities can be described by community-dependent keywords even if they share a semantically similar topic. Articles that are not tagged with eno…

ArticlesInformation RetrievalRetrieval

You are your Metadata: Identification and Obfuscation of Social Media Users using Metadata Information

2018-03-27 · Beatrice Perez, Mirco Musolesi, Gianluca Stringhini

Metadata are associated to most of the information we produce in our daily interactions and communication in the digital world. Yet, surprisingly, metadata are often still catergorized as non-sensitive. Indeed, in the pa…

The CSO Classifier: Ontology-Driven Detection of Research Topics in Scholarly Articles

2021-04-02 · Angelo A. Salatino, Francesco Osborne, Thiviyan Thanapalasingam, Enrico Motta

Classifying research papers according to their research topics is an important task to improve their retrievability, assist the creation of smart analytics, and support a variety of approaches for analysing and making se…

Articles

Elsevier OA CC-By Corpus

2020-08-03 · Daniel Kershaw, Rob Koeling

We introduce the Elsevier OA CC-BY corpus. This is the first open corpus of Scientific Research papers which has a representative sample from across scientific disciplines. This corpus not only includes the full text of …