Sri Lanka Document Datasets: A Large-Scale, Multilingual Resource for Law, News, and Policy
We present a collection of open, machine-readable document datasets covering parliamentary proceedings, legal judgments, government publications, news, and tourism statistics from Sri Lanka. The collection currently comprises of 278,621 documents (80.7 GB) across 26 datasets in Sinhala, Tamil, and English. The datasets are updated daily and mirrored on GitHub and Hugging Face. These resources aim to support research in computational linguistics, legal analytics, socio-political studies, and multilingual natural language processing. We describe the data sources, collection pipeline, formats, and potential use cases, while discussing licensing and ethical considerations. This manuscript is at version v2026-07-02-0940.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Sinhala Language Corpora and Stopwords from a Decade of Sri Lankan Facebook
This paper presents two colloquial Sinhala language corpora from the language efforts of the Data, Analysis and Policy team of LIRNEasia, as well as a list of algorithmically derived stopwords. The larger of the two corp…
A critical examination of crop-yield data for vegetables, maize and Tea for commercialized Sri Lankan biofilm biofertilizers
With increasing global interest in microbial methods for agriculture, the commercialization of biofertilizers in Sri Lanka is of general interest. The use of a biofilm-biofertilizer (BFBF) commercialized in Sri Lanka is …
Trilingual Topic Modeling of Sri Lankan Parliamentary Debates
Sri Lankan parliamentary debates (Hansards) constitute a trilingual corpus of speeches in Sinhala, Tamil, and English, including code-mixed content, yet remain inaccessible to standard NLP pipelines due to layout-complex…
mOSCAR: A Large-scale Multilingual and Multimodal Document-level Corpus
Multimodal Large Language Models (mLLMs) are trained on a large amount of text-image data. While most mLLMs are trained on caption-like data only, Alayrac et al. [2022] showed that additionally training them on interleav…
Few-Shot LearningIn-Context LearningCross-Temporal Sinhala OCR: Page-Level Adaptation and Diachronic Analysis
Sinhala is a morphologically rich abugida spoken by roughly 16 million people in Sri Lanka, and to date, there are no publicly available real-world datasets for page-level Sinhala OCR. All previous studies for assessing …
Document AI