paper-with-me

Papers

Evaluating Legal Reasoning Traces with Legal Issue Tree Rubrics

2025-11-30 · Jinu Lee, Kyoung-Woon On, Simeng Han, Arman Cohan, Julia Hockenmaier arxiv

Evaluating the quality of LLM-generated reasoning traces in expert domains (e.g., law) is essential for ensuring credibility and explainability, yet remains challenging due to the inherent complexity of such reasoning tasks. We introduce LEGIT (LEGal Issue Trees), a novel large-scale (24K instances) expert-level legal reasoning dataset with an emphasis on reasoning trace evaluation. We convert court judgments into hierarchical trees of opposing parties' arguments and the court's conclusions, which serve as rubrics for evaluating the issue coverage and correctness of the reasoning traces. We verify the reliability of these rubrics via human expert annotations and comparison with coarse, less informative rubrics. Using the LEGIT dataset, we show that (1) LLMs' legal reasoning ability is seriously affected by both legal issue coverage and correctness, and that (2) retrieval-augmented generation (RAG) and RL with rubrics bring complementary benefits for legal reasoning abilities, where RAG improves overall reasoning capability, whereas RL improves correctness albeit with reduced coverage.

📄 PDF Abstract BibTeX arXiv:2512.01020

Code (0)

등록된 구현이 없습니다.

Tasks

Legal Reasoning

Similar Papers 제목 키워드 기반

PLawBench: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice

2026-01-23 · Yuzhen Shi, Huanghai Liu, Yiran Hu, Gaojie Song 외 arxiv

As large language models (LLMs) are increasingly applied to legal domain-specific tasks, evaluating their ability to perform legal work in real-world settings has become essential. However, existing legal benchmarks rely…

Legal Reasoning

Evaluating the Role of Large Language Models in Legal Practice in India

2025-08-13 · Rahul Hemrajani arxiv

The integration of Artificial Intelligence(AI) into the legal profession raises significant questions about the capacity of Large Language Models(LLM) to perform key legal tasks. In this paper, I empirically evaluate how…

CALRK-Bench: Evaluating Context-Aware Legal Reasoning in Korean Law

2026-03-27 · JiHyeok Jung, TaeYoung Yoon, HyunSouk Cho arxiv

Legal reasoning requires not only the application of legal rules but also an understanding of the context in which those rules operate. However, existing legal benchmarks primarily evaluate rule application under the ass…

Legal Reasoning

Evaluation of Large Language Models in Legal Applications: Challenges, Methods, and Future Directions

2026-01-21 · Yiran Hu, Huanghai Liu, Chong Wang, Kunran Li 외 arxiv

Large language models (LLMs) are being increasingly integrated into legal applications, including judicial decision support, legal practice assistance, and public-facing legal services. While LLMs show strong potential i…

Legal Reasoning

Evaluating Test-Time Scaling LLMs for Legal Reasoning: OpenAI o1, DeepSeek-R1, and Beyond

2025-03-20 · Yaoyao Yu, Leilei Gan, Yinghao Hu, Bin Wei 외

Recently, Test-Time Scaling Large Language Models (LLMs), such as DeepSeek-R1 and OpenAI o1, have demonstrated exceptional capabilities across various domains and tasks, particularly in reasoning. While these models have…

Legal Reasoning