paper-with-me

홈 › Papers

LLMs Provide Unstable Answers to Legal Questions

2025-01-28 · Andrew Blair-Stanek, Benjamin Van Durme

An LLM is stable if it reaches the same conclusion when asked the identical question multiple times. We find leading LLMs like gpt-4o, claude-3.5, and gemini-1.5 are unstable when providing answers to hard legal questions, even when made as deterministic as possible by setting temperature to 0. We curate and release a novel dataset of 500 legal questions distilled from real cases, involving two parties, with facts, competing legal arguments, and the question of which party should prevail. When provided the exact same question, we observe that LLMs sometimes say one party should win, while other times saying the other party should win. This instability has implications for the increasing numbers of legal AI products, legal processes, and lawyers relying on these LLMs.

📄 PDF Abstract BibTeX arXiv:2502.05196

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FALQU: Finding Answers to Legal Questions

2023-04-12 · Behrooz Mansouri, Ricardo Campos

This paper presents a new test collection for Legal IR, FALQU: Finding Answers to Legal Questions, where questions and answers were obtained from Law Stack Exchange (LawSE), a Q&A website for legal professionals, and oth…

Information RetrievalRetrieval

DeliLaw: A Chinese Legal Counselling System Based on a Large Language Model

2024-08-01 · Nan Xie, Yuelin Bai, Hengyuan Gao, Feiteng Fang 외

Traditional legal retrieval systems designed to retrieve legal documents, statutes, precedents, and other legal information are unable to give satisfactory answers due to lack of semantic understanding of specific questi…

ArticlesHallucinationLanguage ModelingLanguage Modelling+2

LegalBench.PT: A Benchmark for Portuguese Law

2025-02-22 · Beatriz Canaverde, Telmo Pessoa Pires, Leonor Melo Ribeiro, André F. T. Martins

The recent application of LLMs to the legal field has spurred the creation of benchmarks across various jurisdictions and languages. However, no benchmark has yet been specifically designed for the Portuguese legal syste…

Multiple-choice

LEXam: Benchmarking Legal Reasoning on 340 Law Exams

2025-05-19 · Yu Fan, Jingwei Ni, Jakob Merane, Etienne Salimbeni 외

Long-form legal reasoning remains a key challenge for large language models (LLMs) in spite of recent advances in test-time scaling. We introduce LEXam, a novel benchmark derived from 340 law exams spanning 116 law schoo…

BenchmarkingLegal ReasoningMultiple-choice

LeKUBE: A Legal Knowledge Update BEnchmark

2024-07-19 · Changyue Wang, Weihang Su, Hu Yiran, Qingyao Ai 외

Recent advances in Large Language Models (LLMs) have significantly shaped the applications of AI in multiple fields, including the studies of legal intelligence. Trained on extensive legal texts, including statutes and l…

Legal Reasoning