paper-with-me

Papers

Semantic clustering of Russian web search results: possibilities and problems

2014-09-04 · Andrey Kutuzov

The paper deals with word sense induction from lexical co-occurrence graphs. We construct such graphs on large Russian corpora and then apply this data to cluster Mail.ru Search results according to meanings of the query. We compare different methods of performing such clustering and different source corpora. Models of applying distributional semantics to big linguistic data are described.

📄 PDF Abstract BibTeX arXiv:1409.1612

Code (0)

등록된 구현이 없습니다.

Tasks

ClusteringWord Sense Induction

Similar Papers 제목 키워드 기반

RuDSI: graph-based word sense induction dataset for Russian

2022-09-28 · COLING (TextGraphs) 2022 10 · Anna Aksenova, Ekaterina Gavrishina, Elisey Rykov, Andrey Kutuzov

We present RuDSI, a new benchmark for word sense induction (WSI) in Russian. The dataset was created using manual annotation and semi-automatic clustering of Word Usage Graphs (WUGs). Unlike prior WSI datasets for Russia…

ClusteringGraph ClusteringWord Sense Induction

Neural Embedding Language Models in Semantic Clustering of Web Search Results

2016-05-01 · LREC 2016 5 · Andrey Kutuzov, Elizaveta Kuzmenko

In this paper, a new approach towards semantic clustering of the results of ambiguous search queries is presented. We propose using distributed vector representations of words trained with the help of prediction-based ne…

Clustering

Clustering of Russian Adjective-Noun Constructions using Word Embeddings

2017-04-01 · WS 2017 4 · Andrey Kutuzov, Elizaveta Kuzmenko, Lidia Pivovarova

This paper presents a method of automatic construction extraction from a large corpus of Russian. The term {`}construction{'} here means a multi-word expression in which a variable can be replaced with another word from …

ClusteringWord Embeddings

Clustering Comparable Corpora of Russian and Ukrainian Academic Texts: Word Embeddings and Semantic Fingerprints

2016-04-18 · Andrey Kutuzov, Mikhail Kopotev, Tatyana Sviridenko, Lyubov Ivanova

We present our experience in applying distributional semantics (neural word embeddings) to the problem of representing and clustering documents in a bilingual comparable corpus. Our data is a collection of Russian and Uk…

ClusteringTranslationWord Embeddings

Russian News Clustering and Headline Selection Shared Task

2021-05-03 · Ilya Gusev, Ivan Smurov

This paper presents the results of the Russian News Clustering and Headline Selection shared task. As a part of it, we propose the tasks of Russian news event detection, headline selection, and headline generation. These…

ClusteringEvent DetectionHeadline Generation