paper-with-me

홈 › Papers

LegalBench.PT: A Benchmark for Portuguese Law

2025-02-22 · Beatriz Canaverde, Telmo Pessoa Pires, Leonor Melo Ribeiro, André F. T. Martins

The recent application of LLMs to the legal field has spurred the creation of benchmarks across various jurisdictions and languages. However, no benchmark has yet been specifically designed for the Portuguese legal system. In this work, we present LegalBench.PT, the first comprehensive legal benchmark covering key areas of Portuguese law. To develop LegalBench.PT, we first collect long-form questions and answers from real law exams, and then use GPT-4o to convert them into multiple-choice, true/false, and matching formats. Once generated, the questions are filtered and processed to improve the quality of the dataset. To ensure accuracy and relevance, we validate our approach by having a legal professional review a sample of the generated questions. Although the questions are synthetically generated, we show that their basis in human-created exams and our rigorous filtering and processing methods applied result in a reliable benchmark for assessing LLMs' legal knowledge and reasoning abilities. Finally, we evaluate the performance of leading LLMs on LegalBench.PT and investigate potential biases in GPT-4o's responses. We also assess the performance of Portuguese lawyers on a sample of questions to establish a baseline for model comparison and validate the benchmark.

📄 PDF Abstract BibTeX arXiv:2502.16357

Code (0)

등록된 구현이 없습니다.

Tasks

Multiple-choice

Similar Papers 제목 키워드 기반

LegalBench-BR: A Benchmark for Evaluating Large Language Models on Brazilian Legal Decision Classification

2026-04-20 · Pedro Barbosa de Carvalho Neto arxiv

We introduce LegalBench-BR, the first public benchmark for evaluating language models on Brazilian legal text classification. The dataset comprises 3,105 appellate proceedings from the Santa Catarina State Court (TJSC), …

Text Classification

LegalBench-RAG: A Benchmark for Retrieval-Augmented Generation in the Legal Domain

2024-08-19 · Nicholas Pipitone, Ghita Houir Alami

Retrieval-Augmented Generation (RAG) systems are showing promising potential, and are becoming increasingly relevant in AI-powered legal applications. Existing benchmarks, such as LegalBench, assess the generative capabi…

RAGRetrievalRetrieval-augmented Generation

LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

2023-08-20 · NeurIPS 2023 11 · Neel Guha, Julian Nyarko, Daniel E. Ho, Christopher Ré 외

The advent of large language models (LLMs) and their adoption by the legal community has given rise to the question: what types of legal reasoning can LLMs perform? To enable greater study of this question, we present Le…

Legal Reasoning

LegalBench: Prototyping a Collaborative Benchmark for Legal Reasoning

2022-09-13 · Neel Guha, Daniel E. Ho, Julian Nyarko, Christopher Ré

Can foundation models be guided to execute tasks involving legal reasoning? We believe that building a benchmark to answer this question will require sustained collaborative efforts between the computer science and legal…

Legal Reasoning

IslamicLegalBench: Evaluating LLMs Knowledge and Reasoning of Islamic Law Across 1,200 Years of Islamic Pluralist Legal Traditions

2026-02-02 · Ezieddin Elmahjub, Junaid Qadir, Abdullah Mushtaq, Rafay Naeem 외 arxiv

As millions of Muslims turn to LLMs like GPT, Claude, and DeepSeek for religious guidance, a critical question arises: Can these AI systems reliably reason about Islamic law? We introduce IslamicLegalBench, the first ben…

Legal Reasoning