paper-with-me

홈 › Papers

Know thy corpus! Robust methods for digital curation of Web corpora

2020-03-13 · LREC 2020 5 · Serge Sharoff

This paper proposes a novel framework for digital curation of Web corpora in order to provide robust estimation of their parameters, such as their composition and the lexicon. In recent years language models pre-trained on large corpora emerged as clear winners in numerous NLP tasks, but no proper analysis of the corpora which led to their success has been conducted. The paper presents a procedure for robust frequency estimation, which helps in establishing the core lexicon for a given corpus, as well as a procedure for estimating the corpus composition via unsupervised topic models and via supervised genre classification of Web pages. The results of the digital curation study applied to several Web-derived corpora demonstrate their considerable differences. First, this concerns different frequency bursts which impact the core lexicon obtained from each corpus. Second, this concerns the kinds of texts they contain. For example, OpenWebText contains considerably more topical news and political argumentation in comparison to ukWac or Wikipedia. The tools and the results of analysis have been released.

📄 PDF Abstract BibTeX arXiv:2003.06389

Code (1)

ssharoff/robust 공식 구현

Tasks

Genre classificationTopic Models

Similar Papers 제목 키워드 기반

Curatr: A Platform for Semantic Analysis and Curation of Historical Literary Texts

2023-06-13 · Susan Leavy, Gerardine Meaney, Karen Wade, Derek Greene

The increasing availability of digital collections of historical and contemporary literature presents a wealth of possibilities for new research in the humanities. The scale and diversity of such collections however, pre…

DiversityWord Embeddings

Bringing Together Version Control and Quality Assurance of Language Data with LAMA

2022-06-01 · EURALI (LREC) 2022 6 · Aleksandr Riaposov, Elena Lazarenko, Timm Lehmberg

This contribution reports on work in process on project specific software and digital infrastructure components used along with corpus curation workflows in the the framework of the long-term language documentation proje…

Reconstructing NER Corpora: a Case Study on Bulgarian

2020-05-01 · LREC 2020 5 · Iva Marinova, Laska Laskova, Petya Osenova, Kiril Simov 외

The paper reports on the usage of deep learning methods for improving a Named Entity Recognition (NER) training corpus and for predicting and annotating new types in a test corpus. We show how the annotations in a type-b…

Deep Learningnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+1

The ELTE.DH Pilot Corpus -- Creating a Handcrafted Gigaword Web Corpus with Metadata

2020-05-01 · LREC 2020 5 · Bal{\'a}zs Indig, {\'A}rp{\'a}d Knap, Zs{\'o}fia S{\'a}rk{\"o}zi-Lindner, M{\'a}ria Tim{\'a}ri 외

In this article, we present the method we used to create a middle-sized corpus using targeted web crawling. Our corpus contains news portal articles along with their metadata, that can be useful for diverse audiences, ra…

Articles

CALBC: Releasing the Final Corpora

2012-05-01 · LREC 2012 5 · {\c{S}}enay Kafkas, Ian Lewin, David Milward, Erik van Mulligen 외

A number of gold standard corpora for named entity recognition are available to the public. However, the existing gold standard corpora are limited in size and semantic entity types. These usually lead to implementation …

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)