paper-with-me

홈 › Papers

Sri Lanka Document Datasets: A Large-Scale, Multilingual Resource for Law, News, and Policy

2025-10-05 · Nuwan I. Senaratna arxiv

We present a collection of open, machine-readable document datasets covering parliamentary proceedings, legal judgments, government publications, news, and tourism statistics from Sri Lanka. The collection currently comprises of 278,621 documents (80.7 GB) across 26 datasets in Sinhala, Tamil, and English. The datasets are updated daily and mirrored on GitHub and Hugging Face. These resources aim to support research in computational linguistics, legal analytics, socio-political studies, and multilingual natural language processing. We describe the data sources, collection pipeline, formats, and potential use cases, while discussing licensing and ethical considerations. This manuscript is at version v2026-07-02-0940.

📄 PDF Abstract BibTeX arXiv:2510.04124

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Sinhala Language Corpora and Stopwords from a Decade of Sri Lankan Facebook

2020-07-15 · Yudhanjaya Wijeratne, Nisansa de Silva

This paper presents two colloquial Sinhala language corpora from the language efforts of the Data, Analysis and Policy team of LIRNEasia, as well as a list of algorithmically derived stopwords. The larger of the two corp…

A critical examination of crop-yield data for vegetables, maize and Tea for commercialized Sri Lankan biofilm biofertilizers

2023-07-30 · M. W. C. Dharma-wardana, Parakrama Waidyanatha, K. A. Renuka, D. Sumith de S. Abeysiriwardena 외

With increasing global interest in microbial methods for agriculture, the commercialization of biofertilizers in Sri Lanka is of general interest. The use of a biofilm-biofertilizer (BFBF) commercialized in Sri Lanka is …

Trilingual Topic Modeling of Sri Lankan Parliamentary Debates

2026-06-18 · Himath Dhanapala, Haren Daishika, Himandhi Kuruppu, Sithija Seneviratne 외 arxiv

Sri Lankan parliamentary debates (Hansards) constitute a trilingual corpus of speeches in Sinhala, Tamil, and English, including code-mixed content, yet remain inaccessible to standard NLP pipelines due to layout-complex…

mOSCAR: A Large-scale Multilingual and Multimodal Document-level Corpus

2024-06-13 · Matthieu Futeral, Armel Zebaze, Pedro Ortiz Suarez, Julien Abadji 외

Multimodal Large Language Models (mLLMs) are trained on a large amount of text-image data. While most mLLMs are trained on caption-like data only, Alayrac et al. [2022] showed that additionally training them on interleav…

Few-Shot LearningIn-Context Learning

Cross-Temporal Sinhala OCR: Page-Level Adaptation and Diachronic Analysis

2026-06-28 · Avisha Dilhara, Nevidu Jayatilleke arxiv

Sinhala is a morphologically rich abugida spoken by roughly 16 million people in Sri Lanka, and to date, there are no publicly available real-world datasets for page-level Sinhala OCR. All previous studies for assessing …

Document AI