paper-with-me

Papers

One Law, Many Languages: Benchmarking Multilingual Legal Reasoning for Judicial Support

2023-06-15 · Ronja Stern, Vishvaksenan Rasiah, Veton Matoshi, Srinanda Brügger Bose, Matthias Stürmer, Ilias Chalkidis, Daniel E. Ho, Joel Niklaus

Recent strides in Large Language Models (LLMs) have saturated many Natural Language Processing (NLP) benchmarks, emphasizing the need for more challenging ones to properly assess LLM capabilities. However, domain-specific and multilingual benchmarks are rare because they require in-depth expertise to develop. Still, most public models are trained predominantly on English corpora, while other languages remain understudied, particularly for practical domain-specific NLP tasks. In this work, we introduce a novel NLP benchmark for the legal domain that challenges LLMs in five key dimensions: processing \emph{long documents} (up to 50K tokens), using \emph{domain-specific knowledge} (embodied in legal texts), \emph{multilingual} understanding (covering five languages), \emph{multitasking} (comprising legal document-to-document Information Retrieval, Court View Generation, Leading Decision Summarization, Citation Extraction, and eight challenging Text Classification tasks) and \emph{reasoning} (comprising especially Court View Generation, but also the Text Classification tasks). Our benchmark contains diverse datasets from the Swiss legal system, allowing for a comprehensive study of the underlying non-English, inherently multilingual legal system. Despite the large size of our datasets (some with hundreds of thousands of examples), existing publicly available multilingual models struggle with most tasks, even after extensive in-domain pre-training and fine-tuning. We publish all resources (benchmark suite, pre-trained models, code) under permissive open CC BY-SA licenses.

📄 PDF Abstract BibTeX arXiv:2306.09237

Code (2)

stern5497/doc2docbeirir 공식 구현
vr18ub/court_view_generation 공식 구현 pytorch

Tasks

BenchmarkingInformation RetrievalLanguage ModellingLegal Reasoningtext-classificationText Classification

Similar Papers 제목 키워드 기반

Evaluating the Limits of Large Language Models in Multilingual Legal Reasoning

2025-09-26 · Antreas Ioannou, Andreas Shiamishis, Nora Hollenstein, Nezihe Merve Gürel arxiv

In an era dominated by Large Language Models (LLMs), understanding their capabilities and limitations, especially in high-stakes fields like law, is crucial. While LLMs such as Meta's LLaMA, OpenAI's ChatGPT, Google's Ge…

Adversarial RobustnessLegal Reasoning

Bilingual BSARD: Extending Statutory Article Retrieval to Dutch

2024-12-10 · Ehsan Lotfi, Nikolay Banar, Nerses Yuzbashyan, Walter Daelemans

Statutory article retrieval plays a crucial role in making legal information more accessible to both laypeople and legal professionals. Multilingual countries like Belgium present unique challenges for retrieval models d…

ArticlesBenchmarkingRetrieval

MizanQA: Benchmarking Large Language Models on Moroccan Legal Question Answering

2025-08-22 · Adil Bahaj, Mounir Ghogho arxiv

The rapid advancement of large language models (LLMs) has significantly propelled progress in natural language processing (NLP). However, their effectiveness in specialized, low-resource domains-such as Arabic legal cont…

Question AnsweringLegal Reasoning

LEMUR: A Corpus for Robust Fine-Tuning of Multilingual Law Embedding Models for Retrieval

2026-02-10 · Narges Baba Ahmadi, Jan Strich, Martin Semmann, Chris Biemann arxiv

Large language models (LLMs) are increasingly used to access legal information. Yet, their deployment in multilingual legal settings is constrained by unreliable retrieval and the lack of domain-adapted, open-embedding m…

Semantic Retrieval

EUROPA: A Legal Multilingual Keyphrase Generation Dataset

2024-03-01 · Olivier Salaün, Frédéric Piedboeuf, Guillaume Le Berre, David Alfonso Hermelo 외

Keyphrase generation has primarily been explored within the context of academic research articles, with a particular focus on scientific domains and the English language. In this work, we present EUROPA, a dataset for mu…

ArticlesKeyphrase Generation