paper-with-me

홈 › Papers

BERTifying Sinhala -- A Comprehensive Analysis of Pre-trained Language Models for Sinhala Text Classification

2022-08-16 · Vinura Dhananjaya, Piyumal Demotte, Surangika Ranathunga, Sanath Jayasena

This research provides the first comprehensive analysis of the performance of pre-trained language models for Sinhala text classification. We test on a set of different Sinhala text classification tasks and our analysis shows that out of the pre-trained multilingual models that include Sinhala (XLM-R, LaBSE, and LASER), XLM-R is the best model by far for Sinhala text classification. We also pre-train two RoBERTa-based monolingual Sinhala models, which are far superior to the existing pre-trained language models for Sinhala. We show that when fine-tuned, these pre-trained language models set a very strong baseline for Sinhala text classification and are robust in situations where labeled data is insufficient for fine-tuning. We further provide a set of recommendations for using pre-trained models for Sinhala text classification. We also introduce new annotated datasets useful for future research in Sinhala text classification and publicly release our pre-trained models.

📄 PDF Abstract BibTeX arXiv:2208.07864

Code (0)

등록된 구현이 없습니다.

Tasks

Classificationtext-classificationText ClassificationXLM-R

Methods 이 논문이 사용한 방법론

Test 설명 없음
XLM-R XLM-R

Similar Papers 제목 키워드 기반

BERTifying Sinhala - A Comprehensive Analysis of Pre-trained Language Models for Sinhala Text Classification

2022-06-01 · LREC 2022 6 · Vinura Dhananjaya, Piyumal Demotte, Surangika Ranathunga, Sanath Jayasena

This research provides the first comprehensive analysis of the performance of pre-trained language models for Sinhala text classification. We test on a set of different Sinhala text classification tasks and our analysis …

Classificationtext-classificationText ClassificationXLM-R

SiPaKosa: A Comprehensive Corpus of Canonical and Classical Buddhist Texts in Sinhala and Pali

2026-03-31 · Ranidu Gurusinghe, Nevidu Jayatilleke arxiv

SiPaKosa is a comprehensive corpus of Sinhala and Pali doctrinal texts comprising approximately 786K sentences and 9.25M words, incorporating 16 copyright-cleared historical Buddhist documents alongside the complete web-…

Information RetrievalDocument AI

Swa-bhasha Resource Hub: Romanized Sinhala to Sinhala Transliteration Systems and Data Resources

2025-07-12 · Deshan Sumanathilaka, Sameera Perera, Sachithya Dharmasiri, Maneesha Athukorala 외 arxiv

The Swa-bhasha Resource Hub provides a comprehensive collection of data resources and algorithms developed for Romanized Sinhala to Sinhala transliteration between 2020 and 2025. These resources have played a significant…

Sentiment Analysis for Sinhala Language using Deep Learning Techniques

2020-11-14 · Lahiru Senevirathne, Piyumal Demotte, Binod Karunanayake, Udyogi Munasinghe 외

Due to the high impact of the fast-evolving fields of machine learning and deep learning, Natural Language Processing (NLP) tasks have further obtained comprehensive performances for highly resourced languages such as En…

Deep LearningSentiment Analysis

SinSpell: A Comprehensive Spelling Checker for Sinhala

2021-07-07 · Upuli Liyanapathirana, Kaumini Gunasinghe, Gihan Dias

We have built SinSpell, a comprehensive spelling checker for the Sinhala language which is spoken by over 16 million people, mainly in Sri Lanka. However, until recently, Sinhala had no spelling checker with acceptable c…

valid