Clustering of Russian Adjective-Noun Constructions using Word Embeddings
This paper presents a method of automatic construction extraction from a large corpus of Russian. The term {}construction{'} here means a multi-word expression in which a variable can be replaced with another word from the same semantic class, for example, {}a glass of [water/juice/milk]{'}. We deal with constructions that consist of a noun and its adjective modifier. We propose a method of grouping such constructions into semantic classes via 2-step clustering of word vectors in distributional models. We compare it with other clustering techniques and evaluate it against A Russian-English Collocational Dictionary of the Human Body that contains manually annotated groups of constructions with nouns meaning human body parts. The best performing method is used to cluster all adjective-noun bigrams in the Russian National Corpus. Results of this procedure are publicly available and can be used for building Russian construction dictionary as well as to accelerate theoretical studies of constructions.
Code (0)
등록된 구현이 없습니다.
Tasks
ClusteringWord EmbeddingsSimilar Papers 제목 키워드 기반
ConFarm: Extracting Surface Representations of Verb and Noun Constructions from Dependency Annotated Corpora of Russian
ConFarm is a web service dedicated to extraction of surface representations of verb and noun constructions from dependency annotated corpora of Russian texts. Currently, the extraction of constructions with a specific le…
LEMMAELMo and BERT in semantic change detection for Russian
We study the effectiveness of contextualized embeddings for the task of diachronic semantic change detection for Russian language data. Evaluation test sets consist of Russian nouns and adjectives annotated based on thei…
Change DetectionTracing cultural diachronic semantic shifts in Russian using word embeddings: test sets and baselines
The paper introduces manually annotated test sets for the task of tracing diachronic (temporal) semantic shifts in Russian. The two test sets are complementary in that the first one covers comparatively strong semantic c…
General ClassificationWord EmbeddingsExploring the value space of attributes: Unsupervised bidirectional clustering of adjectives in German
The paper presents an iterative bidirectional clustering of adjectives and nouns based on a co-occurrence matrix. The clustering method combines a Vector Space Models (VSM) and the results of a Latent Dirichlet Allocatio…
ClusteringResolving Inflectional Ambiguity of Macedonian Adjectives
Macedonian adjectives are inflected for gender, number, definiteness and degree, with in average 47.98 inflections per headword. The inflection paradigm of qualificative adjectives is even richer, embracing 56.27 morphop…