paper-with-me

Papers

Word Embedding Evaluation for Sinhala

2020-05-01 · LREC 2020 5 · Dimuthu Lakmal, Surangika Ranathunga, Saman Peramuna, Indu Herath

This paper presents the first ever comprehensive evaluation of different types of word embeddings for Sinhala language. Three standard word embedding models, namely, Word2Vec (both Skipgram and CBOW), FastText, and Glove are evaluated under two types of evaluation methods: intrinsic evaluation and extrinsic evaluation. Word analogy and word relatedness evaluations were performed in terms of intrinsic evaluation, while sentiment analysis and part-of-speech (POS) tagging were conducted as the extrinsic evaluation tasks. Benchmark datasets used for intrinsic evaluations were carefully crafted considering specific linguistic features of Sinhala. In general, FastText word embeddings with 300 dimensions reported the finest accuracies across all the evaluation tasks, while Glove reported the lowest results.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Part-Of-Speech TaggingPOSPOS TaggingSentiment AnalysisWord Embeddings

Methods 이 논문이 사용한 방법론

GloVe GloVe Embeddings are a type of word embedding that encode the co-occurrence probability ratio between two words as vector differences. GloVe uses a weighted least squares…
fastText fastText embeddings exploit subword information to construct word embeddings. Representations are learnt of character $n$-grams, and words represented as the sum of the…

Similar Papers 제목 키워드 기반

Sinhala-English Word Embedding Alignment: Introducing Datasets and Benchmark for a Low Resource Language

2023-11-17 · Kasun Wickramasinghe, Nisansa de Silva

Since their inception, embeddings have become a primary ingredient in many flavours of Natural Language Processing (NLP) tasks supplanting earlier types of representation. Even though multilingual embeddings have been us…

Sinhala Sentence Embedding: A Two-Tiered Structure for Low-Resource Languages

2022-10-26 · Gihan Weeraprameshwara, Vihanga Jayawickrama, Nisansa de Silva, Yudhanjaya Wijeratne

In the process of numerically modeling natural languages, developing language embeddings is a vital step. However, it is challenging to develop functional embeddings for resource-poor languages such as Sinhala, for which…

SentenceSentence EmbeddingSentence-EmbeddingSentiment Analysis+2

Keyword Extraction, and Aspect Classification in Sinhala, English, and Code-Mixed Content

2025-04-14 · F. A. Rizvi, T. Navojith, A. M. N. H. Adhikari, W. P. U. Senevirathna 외

Brand reputation in the banking sector is maintained through insightful analysis of customer opinion on code-mixed and multilingual content. Conventional NLP models misclassify or ignore code-mixed text, when mix with lo…

Keyword ExtractionNER

Sinhala Short Sentence Similarity Calculation using Corpus-Based and Knowledge-Based Similarity Measures

2016-12-01 · WS 2016 12 · Jcs Kadupitiya, Surangika Ranathunga, Gihan Dias

Currently, corpus based-similarity, string-based similarity, and knowledge-based similarity techniques are used to compare short phrases. However, no work has been conducted on the similarity of phrases in Sinhala langua…

Semantic SimilaritySemantic Textual SimilaritySentenceSentence Similarity+1

Multi-lingual Mathematical Word Problem Generation using Long Short Term Memory Networks with Enhanced Input Features

2020-05-01 · LREC 2020 5 · Vijini Liyanage, Surangika Ranathunga

A Mathematical Word Problem (MWP) differs from a general textual representation due to the fact that it is comprised of numerical quantities and units, in addition to text. Therefore, MWP generation should be carefully h…

POSTAGWord Embeddings