paper-with-me

홈 › Papers

LEGAL-UQA: A Low-Resource Urdu-English Dataset for Legal Question Answering

2024-10-16 · Faizan Faisal, Umair Yousaf

We present LEGAL-UQA, the first Urdu legal question-answering dataset derived from Pakistan's constitution. This parallel English-Urdu dataset includes 619 question-answer pairs, each with corresponding legal article contexts, addressing the need for domain-specific NLP resources in low-resource languages. We describe the dataset creation process, including OCR extraction, manual refinement, and GPT-4-assisted translation and generation of QA pairs. Our experiments evaluate the latest generalist language and embedding models on LEGAL-UQA, with Claude-3.5-Sonnet achieving 99.19% human-evaluated accuracy. We fine-tune mt5-large-UQA-1.0, highlighting the challenges of adapting multilingual models to specialized domains. Additionally, we assess retrieval performance, finding OpenAI's text-embedding-3-large outperforms Mistral's mistral-embed. LEGAL-UQA bridges the gap between global NLP advancements and localized applications, particularly in constitutional law, and lays the foundation for improved legal information access in Pakistan.

📄 PDF Abstract BibTeX arXiv:2410.13013

Code (1)

nlp-anonymous-researcher/legal-uqa 공식 구현

Tasks

Optical Character Recognition (OCR)Question AnsweringRetrieval

Similar Papers 제목 키워드 기반

MultiLegalPile: A 689GB Multilingual Legal Corpus

2023-06-03 · Joel Niklaus, Veton Matoshi, Matthias Stürmer, Ilias Chalkidis 외

Large, high-quality datasets are crucial for training Large Language Models (LLMs). However, so far, there are few datasets available for specialized critical domains such as law and the available ones are often only for…

MILPaC: A Novel Benchmark for Evaluating Translation of Legal Text to Indian Languages

2023-10-15 · Sayan Mahapatra, Debtanu Datta, Shubham Soni, Adrijit Goswami 외

Most legal text in the Indian judiciary is written in complex English due to historical reasons. However, only a small fraction of the Indian population is comfortable in reading English. Hence legal text needs to be mad…

Machine TranslationTranslation

A Cross-Lingual Statutory Article Retrieval Dataset for Taiwan Legal Studies

2024-10-15 · Yen-Hsiang Wang, Feng-Dian Su, Tzu-Yu Yeh, Yao-Chung Fan

This paper introduces a cross-lingual statutory article retrieval (SAR) dataset designed to enhance legal information retrieval in multilingual settings. Our dataset features spoken-language-style legal inquiries in Engl…

Information RetrievalRetrieval

Retrieval Augmented Generation Framework for the Nepali Legal Domain Question Answering

2026-04-20 · Samir Wagle, Abiral Adhikari, Reewaj Khanal, Batsal Bhandari 외 arxiv

Legal domains in high-resource languages like English have widely adopted artificial intelligence for legal question answering. However, data scarcity in low resource languages such as Nepali has limited the training of …

Question AnsweringAnswer Generation

MILDSum: A Novel Benchmark Dataset for Multilingual Summarization of Indian Legal Case Judgments

2023-10-28 · Debtanu Datta, Shubham Soni, Rajdeep Mukherjee, Saptarshi Ghosh

Automatic summarization of legal case judgments is a practically important problem that has attracted substantial research efforts in many countries. In the context of the Indian judiciary, there is an additional complex…