Zmorge: A German Morphological Lexicon Extracted from Wiktionary
We describe a method to automatically extract a German lexicon from Wiktionary that is compatible with the finite-state morphological grammar SMOR. The main advantage of the resulting lexicon over existing lexica for SMOR is that it is open and permissively licensed. A recall-oriented evaluation shows that a morphological analyser built with our lexicon has comparable coverage compared to existing lexica, and continues to improve as Wiktionary grows. We also describe modifications to the SMOR grammar that result in a more conventional lemmatisation of words.
Code (0)
등록된 구현이 없습니다.
Tasks
Machine TranslationMorphological AnalysisSimilar Papers 제목 키워드 기반
DeLex, a freely-avaible, large-scale and linguistically grounded morphological lexicon for German
We introduce DeLex, a freely-avaible, large-scale and linguistically grounded morphological lexicon for German developed within the Alexina framework. We extracted lexical information from the German wiktionary and devel…
Morphological AnalysisMorphological InflectionENGLAWI: From Human- to Machine-Readable Wiktionary
This paper introduces ENGLAWI, a large, versatile, XML-encoded machine-readable dictionary extracted from Wiktionary. ENGLAWI contains 752,769 articles encoding the full body of information included in Wiktionary: simple…
ArticlesWord EmbeddingsFilling the ___-s in Finnish MWE lexicons
This paper describes the automatic construction of FinnMWE: a lexicon of Finnish Multi-Word Expressions (MWEs). In focus here are syntactic frames: verbal constructions with arguments in a particular morphological form. …
MucLex: A German Lexicon for Surface Realisation
Language resources for languages other than English are often scarce. Rule-based surface realisers need elaborate lexica in order to be able to generate correct language, especially in languages like German, which includ…
Text GenerationUsing Wiktionary to Create Specialized Lexical Resources and Datasets
This paper describes an approach aiming at utilizing Wiktionary data for creating specialized lexical datasets which can be used for enriching other lexical (semantic) resources or for generating datasets that can be use…
Word Sense Disambiguation