paper-with-me

홈 › Papers

Privacy-Preserving Models for Legal Natural Language Processing

2022-11-05 · Ying Yin, Ivan Habernal

Pre-training large transformer models with in-domain data improves domain adaptation and helps gain performance on the domain-specific downstream tasks. However, sharing models pre-trained on potentially sensitive data is prone to adversarial privacy attacks. In this paper, we asked to which extent we can guarantee privacy of pre-training data and, at the same time, achieve better downstream performance on legal tasks without the need of additional labeled data. We extensively experiment with scalable self-supervised learning of transformer models under the formal paradigm of differential privacy and show that under specific training configurations we can improve downstream performance without sacrifying privacy protection for the in-domain data. Our main contribution is utilizing differential privacy for large-scale pre-training of transformer language models in the legal NLP domain, which, to the best of our knowledge, has not been addressed before.

📄 PDF Abstract BibTeX arXiv:2211.02956

Code (1)

trusthlt/privacy-legal-nlp-lm 공식 구현 jax

Tasks

Domain AdaptationPrivacy PreservingSelf-Supervised Learning

Similar Papers 제목 키워드 기반

Towards Task-Agnostic Privacy- and Utility-Preserving Models

2021-09-01 · RANLP 2021 9 · Yaroslav Emelyanov

Modern deep learning models for natural language processing rely heavily on large amounts of annotated texts. However, obtaining such texts may be difficult when they contain personal or confidential information, for exa…

Question Answeringtext-classificationText Classification

Trustworthy AI: Securing Sensitive Data in Large Language Models

2024-09-26 · Georgios Feretzakis, Vassilios S. Verykios

Large Language Models (LLMs) have transformed natural language processing (NLP) by enabling robust text generation and understanding. However, their deployment in sensitive domains like healthcare, finance, and legal ser…

Attributenamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+3

A Short Survey of Viewing Large Language Models in Legal Aspect

2023-03-16 · Zhongxiang Sun

Large language models (LLMs) have transformed many fields, including natural language processing, computer vision, and reinforcement learning. These models have also made a significant impact in the field of law, where t…

Legal Documents Drafting with Fine-Tuned Pre-Trained Large Language Model

2024-06-06 · Chun-Hsien Lin, Pu-Jen Cheng

With the development of large-scale Language Models (LLM), fine-tuning pre-trained LLM has become a mainstream paradigm for solving downstream tasks of natural language processing. However, training a language model in t…

Chinese Word SegmentationLanguage ModelingLanguage ModellingLarge Language Model

GiusBERTo: A Legal Language Model for Personal Data De-identification in Italian Court of Auditors Decisions

2024-06-21 · Giulio Salierno, Rosamaria Bertè, Luca Attias, Carla Morrone 외

Recent advances in Natural Language Processing have demonstrated the effectiveness of pretrained language models like BERT for a variety of downstream tasks. We present GiusBERTo, the first BERT-based model specialized f…

De-identificationLanguage ModelingLanguage Modelling