paper-with-me

Papers

Evaluation of Large Language Models in Legal Applications: Challenges, Methods, and Future Directions

2026-01-21 · Yiran Hu, Huanghai Liu, Chong Wang, Kunran Li, Tien-Hsuan Wu, Haitao Li, Xinran Xu, Siqing Huo, Weihang Su, Ning Zheng, Siyuan Zheng, Qingyao Ai, Yun Liu, Renjun Bian, Yiqun Liu, Charles L. A. Clarke, Weixing Shen, Ben Kao arxiv

Large language models (LLMs) are being increasingly integrated into legal applications, including judicial decision support, legal practice assistance, and public-facing legal services. While LLMs show strong potential in handling legal knowledge and tasks, their deployment in real-world legal settings raises critical concerns beyond surface-level accuracy, involving the soundness of legal reasoning processes and trustworthy issues such as fairness and reliability. Systematic evaluation of LLM performance in legal tasks has therefore become essential for their responsible adoption. This survey identifies key challenges in evaluating LLMs for legal tasks grounded in real-world legal practice. We analyze the major difficulties involved in assessing LLM performance in the legal domain, including outcome correctness, reasoning reliability, and trustworthiness. Building on these challenges, we review and categorize existing evaluation methods and benchmarks according to their task design, datasets, and evaluation metrics. We further discuss the extent to which current approaches address these challenges, highlight their limitations, and outline future research directions toward more realistic, reliable, and legally grounded evaluation frameworks for LLMs in legal domains.

📄 PDF Abstract BibTeX arXiv:2601.15267

Code (0)

등록된 구현이 없습니다.

Tasks

Legal Reasoning

Similar Papers 제목 키워드 기반

LLM Agents in Law: Taxonomy, Applications, and Challenges

2026-01-08 · Shuang Liu, Ruijia Zhang, Ruoyun Ma, Yujia Deng 외 arxiv

Large language models (LLMs) have precipitated a dramatic improvement in the legal domain, yet the deployment of standalone models faces significant limitations regarding hallucination, outdated information, and verifiab…

LexEval: A Comprehensive Chinese Legal Benchmark for Evaluating Large Language Models

2024-09-30 · Haitao Li, You Chen, Qingyao Ai, Yueyue Wu 외

Large language models (LLMs) have made significant progress in natural language processing tasks and demonstrate considerable potential in the legal domain. However, legal applications demand high standards of accuracy, …

Fairness

LeMAJ (Legal LLM-as-a-Judge): Bridging Legal Reasoning and LLM Evaluation

2025-10-08 · Joseph Enguehard, Morgane Van Ermengem, Kate Atkinson, Sujeong Cha 외 arxiv

Evaluating large language model (LLM) outputs in the legal domain presents unique challenges due to the complex and nuanced nature of legal analysis. Current evaluation approaches either depend on reference data, which i…

Legal Reasoning

LeKUBE: A Legal Knowledge Update BEnchmark

2024-07-19 · Changyue Wang, Weihang Su, Hu Yiran, Qingyao Ai 외

Recent advances in Large Language Models (LLMs) have significantly shaped the applications of AI in multiple fields, including the studies of legal intelligence. Trained on extensive legal texts, including statutes and l…

Legal Reasoning

Evaluation Ethics of LLMs in Legal Domain

2024-03-17 · Ruizhe Zhang, Haitao Li, Yueyue Wu, Qingyao Ai 외

In recent years, the utilization of large language models for natural language dialogue has gained momentum, leading to their widespread adoption across various domains. However, their universal competence in addressing …

Ethics