paper-with-me

Papers

SMAuC -- The Scientific Multi-Authorship Corpus

2022-11-04 · Janek Bevendorff, Philipp Sauer, Lukas Gienapp, Wolfgang Kircheis, Erik Körner, Benno Stein, Martin Potthast

The rapidly growing volume of scientific publications offers an interesting challenge for research on methods for analyzing the authorship of documents with one or more authors. However, most existing datasets lack scientific documents or the necessary metadata for constructing new experiments and test cases. We introduce SMAuC, a comprehensive, metadata-rich corpus tailored to scientific authorship analysis. Comprising over 3 million publications across various disciplines from over 5 million authors, SMAuC is the largest openly accessible corpus for this purpose. It encompasses scientific texts from humanities and natural sciences, accompanied by extensive, curated metadata, including unambiguous author IDs. SMAuC aims to significantly advance the domain of authorship analysis in scientific texts.

📄 PDF Abstract BibTeX arXiv:2211.02477

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Quasi Error-free Text Classification and Authorship Recognition in a large Corpus of English Literature based on a Novel Feature Set

2020-10-21 · Arthur M. Jacobs, Annette Kinder

The Gutenberg Literary English Corpus (GLEC) provides a rich source of textual data for research in digital humanities, computational linguistics or neurocognitive poetics. However, so far only a small subcorpus, the Gut…

DiagnosticSentiment Analysistext-classificationText Classification

Entropic selection of concepts unveils hidden topics in documents corpora

2017-05-18 · Andrea Martini, Alessio Cardillo, Paolo De Los Rios

The organization and evolution of science has recently become itself an object of scientific quantitative investigation, thanks to the wealth of information that can be extracted from scientific documents, such as citati…

WEKA in Forensic Authorship Analysis: A corpus-based approach of Saudi Authors

2020-12-01 · ICON 2020 12 · Mashael AlAmr, Eric Atwell

This is a pilot study that aims to explore the potential of using WEKA in forensic authorship analysis. It is a corpus-based research using data from Twitter collected from thirteen authors from Riyadh, Saudi Arabia. It …

CCTAA: A Reproducible Corpus for Chinese Authorship Attribution Research

2022-06-01 · LREC 2022 6 · Haining Wang, Allen Riddell

Authorship attribution infers the likely author of an unsigned, single-authored document from a pool of candidates. Despite recent advances, a lack of standard, reproducible testbeds for Chinese language documents impede…

Authorship Attribution

Research Topic Flows in Co-Authorship Networks

2022-06-16 · Bastian Schäfermeier, Johannes Hirth, Tom Hanika

In scientometrics, scientific collaboration is often analyzed by means of co-authorships. An aspect which is often overlooked and more difficult to quantify is the flow of expertise between authors from different researc…

Community Detection