paper-with-me

홈 › Papers

Sinhala Language Corpora and Stopwords from a Decade of Sri Lankan Facebook

2020-07-15 · Yudhanjaya Wijeratne, Nisansa de Silva

This paper presents two colloquial Sinhala language corpora from the language efforts of the Data, Analysis and Policy team of LIRNEasia, as well as a list of algorithmically derived stopwords. The larger of the two corpora spans 2010 to 2020 and contains 28,825,820 to 29,549,672 words of multilingual text posted by 533 Sri Lankan Facebook pages, including politics, media, celebrities, and other categories; the smaller corpus amounts to 5,402,76 words of only Sinhala text extracted from the larger. Both corpora have markers for their date of creation, page of origin, and content type.

📄 PDF Abstract BibTeX arXiv:2007.07884

Code (1)

LIRNEasia/FacebookDecadeCorpora 공식 구현

Similar Papers 제목 키워드 기반

Sinhala Physical Common Sense Reasoning Dataset for Global PIQA

2026-02-02 · Nisansa de Silva, Surangika Ranathunga arxiv

This paper presents the first-ever Sinhala physical common sense reasoning dataset created as part of Global PIQA. It contains 110 human-created and verified data samples, where each sample consists of a prompt, the corr…

Common Sense Reasoning

Seeking Sinhala Sentiment: Predicting Facebook Reactions of Sinhala Posts

2021-12-01 · Vihanga Jayawickrama, Gihan Weeraprameshwara, Nisansa de Silva, Yudhanjaya Wijeratne

The Facebook network allows its users to record their reactions to text via a typology of emotions. This network, taken at scale, is therefore a prime data set of annotated sentiment data. This paper uses millions of suc…

Binary ClassificationSentiment Analysis

Automatic Creation of a Sentence Aligned Sinhala-Tamil Parallel Corpus

2016-12-01 · WS 2016 12 · Riyafa Abdul Hameed, Nadeeshani Pathirennehelage, Anusha Ihalapathirana, Maryam Ziyad Mohamed 외

A sentence aligned parallel corpus is an important prerequisite in statistical machine translation. However, manual creation of such a parallel corpus is time consuming, and requires experts fluent in both languages. Aut…

Machine TranslationSentenceTranslationWord Alignment

Neural Machine Translation for Sinhala-English Code-Mixed Text

2021-09-01 · RANLP 2021 9 · Archchana Kugathasan, Sagara Sumathipala

Code-mixing has become a moving method of communication among multilingual speakers. Most of the social media content of the multilingual societies are written in code-mixed text. However, most of the current translation…

DecoderMachine TranslationNMTTranslation

SinhalaMMLU: A Comprehensive Benchmark for Evaluating Multitask Language Understanding in Sinhala

2025-09-03 · Ashmari Pramodya, Nirasha Nelki, Heshan Shalinda, Chamila Liyanage 외 arxiv

Large Language Models (LLMs) demonstrate impressive general knowledge and reasoning abilities, yet their evaluation has predominantly focused on global or anglocentric subjects, often neglecting low-resource languages an…

Question AnsweringGeneral Knowledge