paper-with-me

홈 › Papers

LegalCiteBench: Evaluating Citation Reliability in Legal Language Models

2026-05-11 · Sijia Chen, Hang Yin, Shunfan Zhou arxiv

Large language models (LLMs) are increasingly integrated into legal drafting and research workflows, where incorrect citations or fabricated precedents can cause serious professional harm. Existing legal benchmarks largely emphasize statutory reasoning, contract understanding, or general legal question answering, but they do not directly study a central common-law failure mode: when asked to provide case authorities without external grounding, models may return plausible-looking but incorrect citations or cases. We introduce LegalCiteBench, a benchmark for studying closed-book citation recovery, citation verification, and case matching in legal language models. LegalCiteBench contains approximately 24K evaluation instances constructed from 1,000 real U.S. judicial opinions from the Case Law Access Project. The benchmark covers five citation-centric tasks: citation retrieval, citation completion, citation error detection, case matching, and case verification and correction. Across 21 LLMs, exact citation recovery remains highly challenging in this closed-book setting: even the strongest models score below 7/100 on citation retrieval and completion. Within the evaluated models, scale and legal-domain pretraining provide limited gains and do not resolve this difficulty. Models also frequently provide concrete but incorrect or low-overlap authorities under our evaluation protocol, with Misleading Answer Rates (MAR) exceeding 94% for 20 of 21 evaluated models on retrieval-heavy tasks. A prompt-only abstention experiment shows that explicit uncertainty instructions reduce some confident fabrication but do not improve citation correctness. LegalCiteBench is intended as a diagnostic framework for studying authority generation failures, verification behavior, and abstention when external grounding is absent, incomplete, or bypassed.

📄 PDF Abstract BibTeX arXiv:2605.10186

Code (0)

등록된 구현이 없습니다.

Tasks

Question Answering

Similar Papers 제목 키워드 기반

LegalCiteTrust: Benchmarking Citation Trustworthiness in Chinese Long-Form Legal Research Reports

2026-07-23 · Yunhan Li, Mingjie Xie, Zeyang Shi, Gengshen Wu 외 arxiv

Long-form legal research reports increasingly rely on LLMs and agentic research systems, but their reliability depends not only on answering the task, but also on whether cited legal authorities are trustworthy. A citati…

From Judgments to Issues: Structured Extraction of Legal Reasoning with Citation-Hallucination Control

2026-07-03 · Giovanni Piccioli, Alessia Fidelangeli, Piera Santin, Pierpaolo Vivo arxiv

We present an automated pipeline that decomposes Italian tax-court judgments into individual legal issues and extracts, for each issue, a structured XML representation grounded in the IRAC framework and the legal syllogi…

Legal Reasoning

CiteCaseLAW: Citation Worthiness Detection in Caselaw for Legal Assistive Writing

2023-05-03 · Mann Khatri, Pritish Wadhwa, Gitansh Satija, Reshma Sheik 외

In legal document writing, one of the key elements is properly citing the case laws and other sources to substantiate claims and arguments. Understanding the legal domain and identifying appropriate citation context or c…

Citation RecommendationCitation worthinessRecommendation SystemsSpecificity

Incorporating Legal Structure in Retrieval-Augmented Generation: A Case Study on Copyright Fair Use

2025-05-04 · Justin Ho, Alexandra Colby, William Fisher

This paper presents a domain-specific implementation of Retrieval-Augmented Generation (RAG) tailored to the Fair Use Doctrine in U.S. copyright law. Motivated by the increasing prevalence of DMCA takedowns and the lack …

Knowledge GraphsLegal ReasoningRAGRetrieval+1

Chinese Labor Law Large Language Model Benchmark

2026-01-15 · Zixun Lan, Maochun Xu, Yifan Ren, Rui Wu 외 arxiv

Recent advances in large language models (LLMs) have led to substantial progress in domain-specific applications, particularly within the legal domain. However, general-purpose models such as GPT-4 often struggle with sp…

Question Answering