paper-with-me

홈 › Papers

GerDaLIR: A German Dataset for Legal Information Retrieval

2021-11-01 · EMNLP (NLLP) 2021 11 · Marco Wrzalik, Dirk Krechel

We present GerDaLIR, a German Dataset for Legal Information Retrieval based on case documents from the open legal information platform Open Legal Data. The dataset consists of 123K queries, each labelled with at least one relevant document in a collection of 131K case documents. We conduct several baseline experiments including BM25 and a state-of-the-art neural re-ranker. With our dataset, we aim to provide a standardized benchmark for German LIR and promote open research in this area. Beyond that, our dataset comprises sufficient training data to be used as a downstream task for German or multilingual language models.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalRetrieval

Similar Papers 제목 키워드 기반

Finding Needles in Emb(a)dding Haystacks: Legal Document Retrieval via Bagging and SVR Ensembles

2025-01-09 · Kevin Bönisch, Alexander Mehler

We introduce a retrieval approach leveraging Support Vector Regression (SVR) ensembles, bootstrap aggregation (bagging), and embedding spaces on the German Dataset for Legal Information Retrieval (GerDaLIR). By conceptua…

Information RetrievalRetrieval

Chunking German Legal Code

2026-05-19 · Max Prior, Natalia Milanova, Andreas Schultz arxiv

This paper investigates chunking strategies for retrieval-augmented generation on German statutory law, using the German Civil Code as a structured benchmark corpus. We implement and compare a range of segmentation appro…

Computational EfficiencyInformation Retrieval

Segmentation and Processing of German Court Decisions from Open Legal Data

2026-01-04 · Harshil Darji, Martin Heckelmann, Christina Kratsch, Gerard de Melo arxiv

The availability of structured legal data is important for advancing Natural Language Processing (NLP) techniques for the German legal system. One of the most widely used datasets, Open Legal Data, provides a large-scale…

German BERT Model for Legal Named Entity Recognition

2023-03-07 · Harshil Darji, Jelena Mitrović, Michael Granitzer

The use of BERT, one of the most popular language models, has led to improvements in many Natural Language Processing (NLP) tasks. One such task is Named Entity Recognition (NER) i.e. automatic identification of named en…

Language Modellingmodelnamed-entity-recognitionNamed Entity Recognition+4

Summarization of German Court Rulings

2021-11-01 · EMNLP (NLLP) 2021 11 · Ingo Glaser, Sebastian Moser, Florian Matthes

Historically speaking, the German legal language is widely neglected in NLP research, especially in summarization systems, as most of them are based on English newspaper articles. In this paper, we propose the task of au…

Abstractive Text SummarizationArticles