SinhaLegal: A Benchmark Corpus for Information Extraction and Analysis in Sinhala Legislative Texts
SinhaLegal introduces a Sinhala legislative text corpus containing approximately 2 million words across 1,206 legal documents. The dataset includes two types of legal documents: 1,065 Acts dated from 1981 to 2014 and 141 Bills from 2010 to 2014, which were systematically collected from official sources. The texts were extracted using OCR with Google Document AI, followed by extensive post-processing and manual cleaning to ensure high-quality, machine-readable content, along with dedicated metadata files for each document. A comprehensive evaluation was conducted, including corpus statistics, lexical diversity, word frequency analysis, named entity recognition, and topic modelling, demonstrating the structured and domain-specific nature of the corpus. Additionally, perplexity analysis using both large and small language models was performed to assess how effectively language models respond to domain-specific texts. The SinhaLegal corpus represents a vital resource designed to support NLP tasks such as summarisation, information extraction, and analysis, thereby bridging a critical gap in Sinhala legal research.
Code (0)
등록된 구현이 없습니다.
Tasks
Information ExtractionDocument AISimilar Papers 제목 키워드 기반
Span Model for Open Information Extraction on Accurate Corpus
Open information extraction (Open IE) is a challenging task especially due to its brittle data basis. Most of Open IE systems have to be trained on automatically built corpus and evaluated on inaccurate test set. In this…
Open Information ExtractionMethod of Tibetan Person Knowledge Extraction
Person knowledge extraction is the foundation of the Tibetan knowledge graph construction, which provides support for Tibetan question answering system, information retrieval, information extraction and other researches,…
graph constructionInformation RetrievalQuestion AnsweringRetrieval+1\'Etude Exp\'erimentale d'Extraction d'Information dans des Retranscriptions de R\'eunions (An Experimental Approach For Information Extraction in Multi-Party Dialogue Discourse)
Nous nous int{\'e}ressons dans cet article {\`a} l{'}extraction de th{\`e}mes {\`a} partir de retranscriptions textuelles de r{\'e}unions. Ce type de corpus est bruit{\'e}, il manque de formatage, il est peu structur{\'e…
Multi-Source (Pre-)Training for Cross-Domain Measurement, Unit and Context Extraction
We present a cross-domain approach for automated measurement and context extraction based on pre-trained language models. We construct a multi-source, multi-domain corpus and train an end-to-end extraction pipeline. We t…
Domain GeneralizationLarge-Scale Information Extraction from Textual Definitions through Deep Syntactic and Semantic Analysis
We present DefIE, an approach to large-scale Information Extraction (IE) based on a syntactic-semantic analysis of textual definitions. Given a large corpus of definitions we leverage syntactic dependencies to reduce dat…
Open Information ExtractionReading Comprehension