paper-with-me

Papers

Count-Based and Predictive Language Models for Exploring DeReKo

2022-06-01 · CMLC (LREC) 2022 6 · Peter Fankhauser, Marc Kupietz

We present the use of count-based and predictive language models for exploring language use in the German Reference Corpus DeReKo. For collocation analysis along the syntagmatic axis we employ traditional association measures based on co-occurrence counts as well as predictive association measures derived from the output weights of skipgram word embeddings. For inspecting the semantic neighbourhood of words along the paradigmatic axis we visualize the high dimensional word embeddings in two dimensions using t-stochastic neighbourhood embeddings. Together, these visualizations provide a complementary, explorative approach to analysing very large corpora in addition to corpus querying. Moreover, we discuss count-based and predictive models w.r.t. scalability and maintainability in very large corpora.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Word Embeddings

Similar Papers 제목 키워드 기반

Evaluating a Dependency Parser on DeReKo

2020-05-01 · LREC 2020 5 · Peter Fankhauser, Bich-Ngoc Do, Marc Kupietz

We evaluate a graph-based dependency parser on DeReKo, a large corpus of contemporary German. The dependency parser is trained on the German dataset from the SPMRL 2014 Shared Task which contains text from the news domai…

GenitivDB --- a Corpus-Generated Database for German Genitive Classification

2014-05-01 · LREC 2014 5 · Roman Schneider

We present a novel NLP resource for the explanation of linguistic phenomena, built and evaluated exploring very large annotated language corpora. For the compilation, we use the German Reference Corpus (DeReKo) with more…

ClassificationGeneral Classification

Named Entity Tagging a Very Large Unbalanced Corpus: Training and Evaluating NE Classifiers

2014-05-01 · LREC 2014 5 · Joachim Bingel, Thomas Haider

We describe a systematic and application-oriented approach to training and evaluating named entity recognition and classification (NERC) systems, the purpose of which is to identify an optimal system and to train an opti…

ChunkingMachine Translationnamed-entity-recognitionNamed Entity Recognition+3

Recent Developments in DeReKo

2014-05-01 · LREC 2014 5 · Marc Kupietz, Harald L{\"u}ngen

This paper gives an overview of recent developments in the German Reference Corpus DeReKo in terms of growth, maximising relevant corpus strata, metadata, legal issues, and its current and future research interface. Due …

The German Reference Corpus DeReKo: New Developments -- New Opportunities

2018-05-01 · LREC 2018 5 · Marc Kupietz, Harald L{\"u}ngen, Pawe{\l} Kamocki, Andreas Witt
Word Embeddings