paper-with-me

Papers

GenitivDB --- a Corpus-Generated Database for German Genitive Classification

2014-05-01 · LREC 2014 5 · Roman Schneider

We present a novel NLP resource for the explanation of linguistic phenomena, built and evaluated exploring very large annotated language corpora. For the compilation, we use the German Reference Corpus (DeReKo) with more than 5 billion word forms, which is the largest linguistic resource worldwide for the study of contemporary written German. The result is a comprehensive database of German genitive formations, enriched with a broad range of intra- und extralinguistic metadata. It can be used for the notoriously controversial classification and prediction of genitive endings (short endings, long endings, zero-marker). We also evaluate the main factors influencing the use of specific endings. To get a general idea about a factorÂ’s influences and its side effects, we calculate chi-square-tests and visualize the residuals with an association plot. The results are evaluated against a gold standard by implementing tree-based machine learning algorithms. For the statistical analysis, we applied the supervised LMT Logistic Model Trees algorithm, using the WEKA software. We intend to use this gold standard to evaluate GenitivDB, as well as to explore methodologies for a predictive genitive model.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationGeneral Classification

Similar Papers 제목 키워드 기반

Adposition and Case Supersenses v2.6: Guidelines for English

2017-04-07 · Nathan Schneider, Jena D. Hwang, Vivek Srikumar, Archna Bhatia 외

This document offers a detailed linguistic description of SNACS (Semantic Network of Adposition and Case Supersenses; Schneider et al., 2018), an inventory of 52 semantic labels ("supersenses") that characterize the use …

Design and compilation of a specialized Spanish-German parallel corpus

2012-05-01 · LREC 2012 5 · Carla Parra Escart{\'\i}n

This paper discusses the design and compilation of the TRIS corpus, a specialized parallel corpus of Spanish and German texts. It will be used for phraseological research aimed at improving statistical machine translatio…

Machine TranslationSentenceTranslation

German \textitnach-Particle Verbs in Semantic Theory and Corpus Data

2012-05-01 · LREC 2012 5 · Boris Haselbach, Wolfgang Seeker, Kerstin Eckart

In this paper, we present a database-supported corpus study where we combine automatically obtained linguistic information from a statistical dependency parser, namely the occurrence of a dative argument, with prediction…

Management

Augmenting a German Morphological Database by Data-Intense Methods

2019-08-01 · WS 2019 8 · Petra Steiner

This paper deals with the automatic enhancement of a new German morphological database. While there are some databases for flat word segmentation, this is the first available resource which can be directly used for deep …

LEMMA

Designing a Bilingual Speech Corpus for French and German Language Learners: a Two-Step Process

2014-05-01 · LREC 2014 5 · Camille Fauth, Anne Bonneau, Frank Zimmerer, Juergen Trouvain 외

We present the design of a corpus of native and non-native speech for the language pair French-German, with a special emphasis on phonetic and prosodic aspects. To our knowledge there is no suitable corpus, in terms of s…