GhoSt-NN: A Representative Gold Standard of German Noun-Noun Compounds
This paper presents a novel gold standard of German noun-noun compounds (Ghost-NN) including 868 compounds annotated with corpus frequencies of the compounds and their constituents, productivity and ambiguity of the constituents, semantic relations between the constituents, and compositionality ratings of compound-constituent pairs. Moreover, a subset of the compounds containing 180 compounds is balanced for the productivity of the modifiers (distinguishing low/mid/high productivity) and the ambiguity of the heads (distinguishing between heads with 1, 2 and {\textgreater}2 senses
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
GhoSt-PV: A Representative Gold Standard of German Particle Verbs
German particle verbs represent a frequent type of multi-word-expression that forms a highly productive paradigm in the lexicon. Similarly to other multi-word expressions, particle verbs exhibit various levels of composi…
Animacy Denoting German Nouns: Annotation and Classification
In this paper, we introduce a gold standard for animacy detection comprising almost 14,500 German nouns that might be used to denote either animate entities or non-animate entities. We present inter-annotator agreement o…
ClassificationPolar Quantification of Actor Noun Phrases for German
In this paper, we discuss work that strives to measure the degree of negativity - the negative polar load - of noun phrases, especially those denoting actors. Since no gold standard data is available for German for this …
All That Glitters is Not Gold: A Gold Standard of Adjective-Noun Collocations for German
In this paper we present the GerCo dataset of adjective-noun collocations for German, such as alter Freund {`}old friend{'} and tiefe Liebe {`}deep love{'}. The annotation has been performed by experts based on the annot…
AllWord EmbeddingsProbing BERT for German Compound Semantics
This paper investigates the extent to which pretrained German BERT encodes knowledge of noun compound semantics. We comprehensively vary combinations of target tokens, layers, and cased vs. uncased models, and evaluate t…