Recent Developments in DeReKo
This paper gives an overview of recent developments in the German Reference Corpus DeReKo in terms of growth, maximising relevant corpus strata, metadata, legal issues, and its current and future research interface. Due to the recent acquisition of new licenses, DeReKo has grown by a factor of four in the first half of 2014, mostly in the area of newspaper text, and presently contains over 24 billion word tokens. Other strata, like fictional texts, web corpora, in particular CMC texts, and spoken but conceptually written texts have also increased significantly. We report on the newly acquired corpora that led to the major increase, on the principles and strategies behind our corpus acquisition activities, and on our solutions for the emerging legal, organisational, and technical challenges.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
The German Reference Corpus DeReKo: New Developments -- New Opportunities
Evaluating a Dependency Parser on DeReKo
We evaluate a graph-based dependency parser on DeReKo, a large corpus of contemporary German. The dependency parser is trained on the German dataset from the SPMRL 2014 Shared Task which contains text from the news domai…
Named Entity Tagging a Very Large Unbalanced Corpus: Training and Evaluating NE Classifiers
We describe a systematic and application-oriented approach to training and evaluating named entity recognition and classification (NERC) systems, the purpose of which is to identify an optimal system and to train an opti…
ChunkingMachine Translationnamed-entity-recognitionNamed Entity Recognition+3RKorAPClient: An R Package for Accessing the German Reference Corpus DeReKo via KorAP
Making corpora accessible and usable for linguistic research is a huge challenge in view of (too) big data, legal issues and a rapidly evolving methodology. This does not only affect the design of user-friendly graphical…
Count-Based and Predictive Language Models for Exploring DeReKo
We present the use of count-based and predictive language models for exploring language use in the German Reference Corpus DeReKo. For collocation analysis along the syntagmatic axis we employ traditional association mea…
Word Embeddings