paper-with-me

Papers

SinhaLegal: A Benchmark Corpus for Information Extraction and Analysis in Sinhala Legislative Texts

2026-03-05 · Minduli Lasandi, Nevidu Jayatilleke arxiv

SinhaLegal introduces a Sinhala legislative text corpus containing approximately 2 million words across 1,206 legal documents. The dataset includes two types of legal documents: 1,065 Acts dated from 1981 to 2014 and 141 Bills from 2010 to 2014, which were systematically collected from official sources. The texts were extracted using OCR with Google Document AI, followed by extensive post-processing and manual cleaning to ensure high-quality, machine-readable content, along with dedicated metadata files for each document. A comprehensive evaluation was conducted, including corpus statistics, lexical diversity, word frequency analysis, named entity recognition, and topic modelling, demonstrating the structured and domain-specific nature of the corpus. Additionally, perplexity analysis using both large and small language models was performed to assess how effectively language models respond to domain-specific texts. The SinhaLegal corpus represents a vital resource designed to support NLP tasks such as summarisation, information extraction, and analysis, thereby bridging a critical gap in Sinhala legal research.

📄 PDF Abstract BibTeX arXiv:2603.04854

Code (0)

등록된 구현이 없습니다.

Tasks

Information ExtractionDocument AI

Similar Papers 제목 키워드 기반

Span Model for Open Information Extraction on Accurate Corpus

2019-01-30 · Junlang Zhan, Hai Zhao

Open information extraction (Open IE) is a challenging task especially due to its brittle data basis. Most of Open IE systems have to be trained on automatically built corpus and evaluated on inaccurate test set. In this…

Open Information Extraction

Method of Tibetan Person Knowledge Extraction

2016-04-11 · Yuan Sun, Zhen Zhu

Person knowledge extraction is the foundation of the Tibetan knowledge graph construction, which provides support for Tibetan question answering system, information retrieval, information extraction and other researches,…

graph constructionInformation RetrievalQuestion AnsweringRetrieval+1

\'Etude Exp\'erimentale d'Extraction d'Information dans des Retranscriptions de R\'eunions (An Experimental Approach For Information Extraction in Multi-Party Dialogue Discourse)

2018-05-01 · JEPTALNRECITAL 2018 5 · Pegah Alizadeh, Peggy Cellier, Thierry Charnois, Bruno Cremilleux 외

Nous nous int{\'e}ressons dans cet article {\`a} l{'}extraction de th{\`e}mes {\`a} partir de retranscriptions textuelles de r{\'e}unions. Ce type de corpus est bruit{\'e}, il manque de formatage, il est peu structur{\'e…

Multi-Source (Pre-)Training for Cross-Domain Measurement, Unit and Context Extraction

2023-08-05 · Yueling Li, Sebastian Martschat, Simone Paolo Ponzetto

We present a cross-domain approach for automated measurement and context extraction based on pre-trained language models. We construct a multi-source, multi-domain corpus and train an end-to-end extraction pipeline. We t…

Domain Generalization

Large-Scale Information Extraction from Textual Definitions through Deep Syntactic and Semantic Analysis

2015-01-01 · TACL 2015 1 · Claudio Delli Bovi, Luca Telesca, Roberto Navigli

We present DefIE, an approach to large-scale Information Extraction (IE) based on a syntactic-semantic analysis of textual definitions. Given a large corpus of definitions we leverage syntactic dependencies to reduce dat…

Open Information ExtractionReading Comprehension