paper-with-me

홈 › Papers

Sinhala Short Sentence Similarity Calculation using Corpus-Based and Knowledge-Based Similarity Measures

2016-12-01 · WS 2016 12 · Jcs Kadupitiya, Surangika Ranathunga, Gihan Dias

Currently, corpus based-similarity, string-based similarity, and knowledge-based similarity techniques are used to compare short phrases. However, no work has been conducted on the similarity of phrases in Sinhala language. In this paper, we present a hybrid methodology to compute the similarity between two Sinhala sentences using a Semantic Similarity Measurement technique (corpus-based similarity measurement plus knowledge-based similarity measurement) that makes use of word order information. Since Sinhala WordNet is still under construction, we used lexical resources in performing this semantic similarity calculation. Evaluation using 4000 sentence pairs yielded an average MSE of 0.145 and a Pearson correla-tion factor of 0.832.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic SimilaritySemantic Textual SimilaritySentenceSentence SimilarityWord Sense Disambiguation

Similar Papers 제목 키워드 기반

Automatic Creation of a Sentence Aligned Sinhala-Tamil Parallel Corpus

2016-12-01 · WS 2016 12 · Riyafa Abdul Hameed, Nadeeshani Pathirennehelage, Anusha Ihalapathirana, Maryam Ziyad Mohamed 외

A sentence aligned parallel corpus is an important prerequisite in statistical machine translation. However, manual creation of such a parallel corpus is time consuming, and requires experts fluent in both languages. Aut…

Machine TranslationSentenceTranslationWord Alignment

Neural Machine Translation for Sinhala-English Code-Mixed Text

2021-09-01 · RANLP 2021 9 · Archchana Kugathasan, Sagara Sumathipala

Code-mixing has become a moving method of communication among multilingual speakers. Most of the social media content of the multilingual societies are written in code-mixed text. However, most of the current translation…

DecoderMachine TranslationNMTTranslation

SiPaKosa: A Comprehensive Corpus of Canonical and Classical Buddhist Texts in Sinhala and Pali

2026-03-31 · Ranidu Gurusinghe, Nevidu Jayatilleke arxiv

SiPaKosa is a comprehensive corpus of Sinhala and Pali doctrinal texts comprising approximately 786K sentences and 9.25M words, incorporating 16 copyright-cleared historical Buddhist documents alongside the complete web-…

Information RetrievalDocument AI

HelaBERT: Enhancing Sinhala Language Understanding with Dual Pooling Classification Head

2026-08-24 · Thisen Ekanayake, Nisansa de Silva arxiv

We present HelaBERT, a family of two BERT-based masked language models pre-trained from scratch on approximately 1 billion tokens of Sinhala text sourced from MADLAD-400, CulturaX, and a custom corpus comprising news art…

Text ClassificationSentiment Analysis

Adapting the Tesseract Open-Source OCR Engine for Tamil and Sinhala Legacy Fonts and Creating a Parallel Corpus for Tamil-Sinhala-English

2021-09-13 · Charangan Vasantharajan, Laksika Tharmalingam, Uthayasanker Thayasivam

Most low-resource languages do not have the necessary resources to create even a substantial monolingual corpus. These languages may often be found in government proceedings but mainly in Portable Document Format (PDF) t…

Optical Character Recognition (OCR)