paper-with-me

Papers

Spanish Datasets for Sensitive Entity Detection in the Legal Domain

2022-06-01 · LREC 2022 6 · Ona de Gibert Bonet, Aitor García Pablos, Montse Cuadros, Maite Melero

The de-identification of sensible data, also known as automatic textual anonymisation, is essential for data sharing and reuse, both for research and commercial purposes. The first step for data anonymisation is the detection of sensible entities. In this work, we present four new datasets for named entity detection in Spanish in the legal domain. These datasets have been generated in the framework of the MAPA project, three smaller datasets have been manually annotated and one large dataset has been automatically annotated, with an estimated error rate of around 14%. In order to assess the quality of the generated datasets, we have used them to fine-tune a battery of entity-detection models, using as foundation different pre-trained language models: one multilingual, two general-domain monolingual and one in-domain monolingual. We compare the results obtained, which validate the datasets as a valuable resource to fine-tune models for the task of named entity detection. We further explore the proposed methodology by applying it to a real use case scenario.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

De-identification

Similar Papers 제목 키워드 기반

Legal-ES: A Set of Large Scale Resources for Spanish Legal Text Processing

2020-05-01 · LREC 2020 5 · Doaa Samy, Jer{\'o}nimo Arenas-Garc{\'\i}a, David P{\'e}rez-Fern{\'a}ndez

Legal-ES is an open source resource kit for legal Spanish. It consists of a large scale Spanish corpus of open legal texts and different kinds of language models including word embeddings and topic models. The corpus inc…

NavigateSemantic SimilaritySemantic Textual SimilarityTopic Models+1

Spanish Legalese Language Model and Corpora

2021-10-23 · Asier Gutiérrez-Fandiño, Jordi Armengol-Estapé, Aitor Gonzalez-Agirre, Marta Villegas

There are many Language Models for the English language according to its worldwide relevance. However, for the Spanish language, even if it is a widely spoken language, there are very few Spanish Language Models which re…

Language ModelingLanguage Modellingmodel

LegalGuardian: A Privacy-Preserving Framework for Secure Integration of Large Language Models in Legal Practice

2025-01-19 · M. Mikail Demir, Hakan T. Otal, M. Abdullah Canbaz

Large Language Models (LLMs) hold promise for advancing legal practice by automating complex tasks and improving access to justice. However, their adoption is limited by concerns over client confidentiality, especially w…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+3

3CEL: A corpus of legal Spanish contract clauses

2025-01-27 · Nuria Aldama García, Patricia Marsà Morales, David Betancur Sánchez, Álvaro Barbero Jiménez 외

Legal corpora for Natural Language Processing (NLP) are valuable and scarce resources in languages like Spanish due to two main reasons: data accessibility and legal expert knowledge availability. INESData 2024 is a Euro…

MEL: Legal Spanish Language Model

2025-01-27 · David Betancur Sánchez, Nuria Aldama García, Álvaro Barbero Jiménez, Marta Guerrero Nieto 외

Legal texts, characterized by complex and specialized terminology, present a significant challenge for Language Models. Adding an underrepresented language, such as Spanish, to the mix makes it even more challenging. Whi…

Language ModelingLanguage Modellingmodel