paper-with-me

Papers

FairLex: A Multilingual Benchmark for Evaluating Fairness in Legal Text Processing

2022-03-14 · ACL 2022 5 · Ilias Chalkidis, Tommaso Pasini, Sheng Zhang, Letizia Tomada, Sebastian Felix Schwemer, Anders Søgaard

We present a benchmark suite of four datasets for evaluating the fairness of pre-trained language models and the techniques used to fine-tune them for downstream tasks. Our benchmarks cover four jurisdictions (European Council, USA, Switzerland, and China), five languages (English, German, French, Italian and Chinese) and fairness across five attributes (gender, age, region, language, and legal area). In our experiments, we evaluate pre-trained language models using several group-robust fine-tuning techniques and show that performance group disparities are vibrant in many cases, while none of these techniques guarantee fairness, nor consistently mitigate group disparities. Furthermore, we provide a quantitative and qualitative analysis of our results, highlighting open challenges in the development of robustness methods in legal NLP.

📄 PDF Abstract BibTeX arXiv:2203.07228

Code (1)

coastalcph/fairlex 공식 구현 pytorch

Tasks

Fairness

Similar Papers 제목 키워드 기반

FairLex: A Multilingual Benchmark for Evaluating Fairness in Legal Text Processing

2021-11-16 · ACL ARR November 2021 11 · Anonymous

We present a benchmark suite of four datasets for evaluating the fairness of pre-trained legal language models and the techniques used to fine-tune them for downstream tasks. Our benchmarks cover four jurisdictions (Euro…

Fairness

Towards Explainability and Fairness in Swiss Judgement Prediction: Benchmarking on a Multilingual Dataset

2024-02-26 · Santosh T. Y. S. S, Nina Baumgartner, Matthias Stürmer, Matthias Grabmair 외

The assessment of explainability in Legal Judgement Prediction (LJP) systems is of paramount importance in building trustworthy and transparent systems, particularly considering the reliance of these systems on factors t…

BenchmarkingCross-Lingual TransferData AugmentationFairness+1

When Fairness Isn't Statistical: The Limits of Machine Learning in Evaluating Legal Reasoning

2025-06-04 · Claire Barale, Michael Rovatsos, Nehal Bhuta

Legal decisions are increasingly evaluated for fairness, consistency, and bias using machine learning (ML) techniques. In high-stakes domains like refugee adjudication, such methods are often applied to detect disparitie…

ClusteringFairnessLegal Reasoning

Evaluating the Limits of Large Language Models in Multilingual Legal Reasoning

2025-09-26 · Antreas Ioannou, Andreas Shiamishis, Nora Hollenstein, Nezihe Merve Gürel arxiv

In an era dominated by Large Language Models (LLMs), understanding their capabilities and limitations, especially in high-stakes fields like law, is crucial. While LLMs such as Meta's LLaMA, OpenAI's ChatGPT, Google's Ge…

Adversarial RobustnessLegal Reasoning

JustEva: A Toolkit to Evaluate LLM Fairness in Legal Knowledge Inference

2025-09-15 · Zongyue Xue, Siyuan Zheng, Shaochun Wang, Yiran Hu 외 arxiv

The integration of Large Language Models (LLMs) into legal practice raises pressing concerns about judicial fairness, particularly due to the nature of their "black-box" processes. This study introduces JustEva, a compre…