Evaluation of Croatian Word Embeddings
Croatian is poorly resourced and highly inflected language from Slavic language family. Nowadays, research is focusing mostly on English. We created a new word analogy corpus based on the original English Word2vec word analogy corpus and added some of the specific linguistic aspects from Croatian language. Next, we created Croatian WordSim353 and RG65 corpora for a basic evaluation of word similarities. We compared created corpora on two popular word representation models, based on Word2Vec tool and fastText tool. Models has been trained on 1.37B tokens training data corpus and tested on a new robust Croatian word analogy corpus. Results show that models are able to create meaningful word representation. This research has shown that free word order and the higher morphological complexity of Croatian language influences the quality of resulting word embeddings.
Code (1)
Tasks
Word EmbeddingsMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Comparison of Short-Text Sentiment Analysis Methods for Croatian
We focus on the task of supervised sentiment classification of short and informal texts in Croatian, using two simple yet effective methods: word embeddings and string kernels. We investigate whether word embeddings offe…
General ClassificationSentiment AnalysisSentiment ClassificationStock Price Prediction+2Characterizing Linguistic Shifts in Croatian News via Diachronic Word Embeddings
Measuring how semantics of words change over time improves our understanding of how cultures and perspectives change. Diachronic word embeddings help us quantify this shift, although previous studies leveraged substantia…
ArticlesDiachronic Word EmbeddingsSentiment AnalysisWord EmbeddingsMultilingual Culture-Independent Word Analogy Datasets
In text processing, deep neural networks mostly use word embeddings as an input. Embeddings have to ensure that relations between words are reflected through distances in a high-dimensional numeric space. To compare the …
Cultural Vocal Bursts Intensity PredictionWord EmbeddingsA bilingual approach to specialised adjectives through word embeddings in the karstology domain
We present an experiment in extracting adjectives which express a specific semantic relation using word embeddings. The results of the experiment are then thoroughly analysed and categorised into groups of adjectives exh…
Semantic SimilaritySemantic Textual SimilarityWord EmbeddingsGraph-Based Induction of Word Senses in Croatian
Word sense induction (WSI) seeks to induce senses of words from unannotated corpora. In this paper, we address the WSI task for the Croatian language. We adopt the word clustering approach based on co-occurrence graphs, …
Clusteringgraph constructionWord Sense DisambiguationWord Sense Induction