Word Embedding Evaluation for Sinhala
This paper presents the first ever comprehensive evaluation of different types of word embeddings for Sinhala language. Three standard word embedding models, namely, Word2Vec (both Skipgram and CBOW), FastText, and Glove are evaluated under two types of evaluation methods: intrinsic evaluation and extrinsic evaluation. Word analogy and word relatedness evaluations were performed in terms of intrinsic evaluation, while sentiment analysis and part-of-speech (POS) tagging were conducted as the extrinsic evaluation tasks. Benchmark datasets used for intrinsic evaluations were carefully crafted considering specific linguistic features of Sinhala. In general, FastText word embeddings with 300 dimensions reported the finest accuracies across all the evaluation tasks, while Glove reported the lowest results.
Code (0)
등록된 구현이 없습니다.
Tasks
Part-Of-Speech TaggingPOSPOS TaggingSentiment AnalysisWord EmbeddingsMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Sinhala-English Word Embedding Alignment: Introducing Datasets and Benchmark for a Low Resource Language
Since their inception, embeddings have become a primary ingredient in many flavours of Natural Language Processing (NLP) tasks supplanting earlier types of representation. Even though multilingual embeddings have been us…
Sinhala Sentence Embedding: A Two-Tiered Structure for Low-Resource Languages
In the process of numerically modeling natural languages, developing language embeddings is a vital step. However, it is challenging to develop functional embeddings for resource-poor languages such as Sinhala, for which…
SentenceSentence EmbeddingSentence-EmbeddingSentiment Analysis+2Keyword Extraction, and Aspect Classification in Sinhala, English, and Code-Mixed Content
Brand reputation in the banking sector is maintained through insightful analysis of customer opinion on code-mixed and multilingual content. Conventional NLP models misclassify or ignore code-mixed text, when mix with lo…
Keyword ExtractionNERSinhala Short Sentence Similarity Calculation using Corpus-Based and Knowledge-Based Similarity Measures
Currently, corpus based-similarity, string-based similarity, and knowledge-based similarity techniques are used to compare short phrases. However, no work has been conducted on the similarity of phrases in Sinhala langua…
Semantic SimilaritySemantic Textual SimilaritySentenceSentence Similarity+1Multi-lingual Mathematical Word Problem Generation using Long Short Term Memory Networks with Enhanced Input Features
A Mathematical Word Problem (MWP) differs from a general textual representation due to the fact that it is comprised of numerical quantities and units, in addition to text. Therefore, MWP generation should be carefully h…
POSTAGWord Embeddings