paper-with-me

홈 › Papers

Survey on Publicly Available Sinhala Natural Language Processing Tools and Research

2019-06-05 · Nisansa de Silva

Sinhala is the native language of the Sinhalese people who make up the largest ethnic group of Sri Lanka. The language belongs to the globe-spanning language tree, Indo-European. However, due to poverty in both linguistic and economic capital, Sinhala, in the perspective of Natural Language Processing tools and research, remains a resource-poor language which has neither the economic drive its cousin English has nor the sheer push of the law of numbers a language such as Chinese has. A number of research groups from Sri Lanka have noticed this dearth and the resultant dire need for proper tools and research for Sinhala natural language processing. However, due to various reasons, these attempts seem to lack coordination and awareness of each other. The objective of this paper is to fill that gap of a comprehensive literature survey of the publicly available Sinhala natural language tools and research so that the researchers working in this field can better utilize contributions of their peers. As such, we shall be uploading this paper to arXiv and perpetually update it periodically to reflect the advances made in the field.

📄 PDF Abstract BibTeX arXiv:1906.02358

Code (1)

lknlp/lknlp.github.io 공식 구현

Similar Papers 제목 키워드 기반

Swa-bhasha Resource Hub: Romanized Sinhala to Sinhala Transliteration Systems and Data Resources

2025-07-12 · Deshan Sumanathilaka, Sameera Perera, Sachithya Dharmasiri, Maneesha Athukorala 외 arxiv

The Swa-bhasha Resource Hub provides a comprehensive collection of data resources and algorithms developed for Romanized Sinhala to Sinhala transliteration between 2020 and 2025. These resources have played a significant…

SOLD: Sinhala Offensive Language Dataset

2022-12-01 · Tharindu Ranasinghe, Isuri Anuradha, Damith Premasiri, Kanishka Silva 외

The widespread of offensive content online, such as hate speech and cyber-bullying, is a global phenomenon. This has sparked interest in the artificial intelligence (AI) and natural language processing (NLP) communities,…

Language IdentificationSentence

A Systematic Approach to Derive a Refined Speech Corpus for Sinhala

2022-06-01 · LREC 2022 6 · Disura Warusawithana, Nilmani Kulaweera, Lakshan Weerasinghe, Buddhika Karunarathne

Speech Recognition is an active research area where advances of technology have continuously driven the development of research work. However, due to the lack of adequate resources, certain languages such as Sinhala, are…

speech-recognitionSpeech Recognition

NSINA: A News Corpus for Sinhala

2024-03-25 · Hansi Hettiarachchi, Damith Premasiri, Lasitha Uyangodage, Tharindu Ranasinghe

The introduction of large language models (LLMs) has advanced natural language processing (NLP), but their effectiveness is largely dependent on pre-training resources. This is especially evident in low-resource language…

ArticlesBenchmarkingHeadline Generation

BERTifying Sinhala - A Comprehensive Analysis of Pre-trained Language Models for Sinhala Text Classification

2022-06-01 · LREC 2022 6 · Vinura Dhananjaya, Piyumal Demotte, Surangika Ranathunga, Sanath Jayasena

This research provides the first comprehensive analysis of the performance of pre-trained language models for Sinhala text classification. We test on a set of different Sinhala text classification tasks and our analysis …

Classificationtext-classificationText ClassificationXLM-R