paper-with-me

Papers

Recent Developments in DeReKo

2014-05-01 · LREC 2014 5 · Marc Kupietz, Harald L{\"u}ngen

This paper gives an overview of recent developments in the German Reference Corpus DeReKo in terms of growth, maximising relevant corpus strata, metadata, legal issues, and its current and future research interface. Due to the recent acquisition of new licenses, DeReKo has grown by a factor of four in the first half of 2014, mostly in the area of newspaper text, and presently contains over 24 billion word tokens. Other strata, like fictional texts, web corpora, in particular CMC texts, and spoken but conceptually written texts have also increased significantly. We report on the newly acquired corpora that led to the major increase, on the principles and strategies behind our corpus acquisition activities, and on our solutions for the emerging legal, organisational, and technical challenges.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The German Reference Corpus DeReKo: New Developments -- New Opportunities

2018-05-01 · LREC 2018 5 · Marc Kupietz, Harald L{\"u}ngen, Pawe{\l} Kamocki, Andreas Witt
Word Embeddings

Evaluating a Dependency Parser on DeReKo

2020-05-01 · LREC 2020 5 · Peter Fankhauser, Bich-Ngoc Do, Marc Kupietz

We evaluate a graph-based dependency parser on DeReKo, a large corpus of contemporary German. The dependency parser is trained on the German dataset from the SPMRL 2014 Shared Task which contains text from the news domai…

Named Entity Tagging a Very Large Unbalanced Corpus: Training and Evaluating NE Classifiers

2014-05-01 · LREC 2014 5 · Joachim Bingel, Thomas Haider

We describe a systematic and application-oriented approach to training and evaluating named entity recognition and classification (NERC) systems, the purpose of which is to identify an optimal system and to train an opti…

ChunkingMachine Translationnamed-entity-recognitionNamed Entity Recognition+3

RKorAPClient: An R Package for Accessing the German Reference Corpus DeReKo via KorAP

2020-05-01 · LREC 2020 5 · Marc Kupietz, Nils Diewald, Eliza Margaretha

Making corpora accessible and usable for linguistic research is a huge challenge in view of (too) big data, legal issues and a rapidly evolving methodology. This does not only affect the design of user-friendly graphical…

Count-Based and Predictive Language Models for Exploring DeReKo

2022-06-01 · CMLC (LREC) 2022 6 · Peter Fankhauser, Marc Kupietz

We present the use of count-based and predictive language models for exploring language use in the German Reference Corpus DeReKo. For collocation analysis along the syntagmatic axis we employ traditional association mea…

Word Embeddings