paper-with-me

Papers

LeKUBE: A Legal Knowledge Update BEnchmark

2024-07-19 · Changyue Wang, Weihang Su, Hu Yiran, Qingyao Ai, Yueyue Wu, Cheng Luo, Yiqun Liu, Min Zhang, Shaoping Ma

Recent advances in Large Language Models (LLMs) have significantly shaped the applications of AI in multiple fields, including the studies of legal intelligence. Trained on extensive legal texts, including statutes and legal documents, the legal LLMs can capture important legal knowledge/concepts effectively and provide important support for downstream legal applications such as legal consultancy. Yet, the dynamic nature of legal statutes and interpretations also poses new challenges to the use of LLMs in legal applications. Particularly, how to update the legal knowledge of LLMs effectively and efficiently has become an important research problem in practice. Existing benchmarks for evaluating knowledge update methods are mostly designed for the open domain and cannot address the specific challenges of the legal domain, such as the nuanced application of new legal knowledge, the complexity and lengthiness of legal regulations, and the intricate nature of legal reasoning. To address this gap, we introduce the Legal Knowledge Update BEnchmark, i.e. LeKUBE, which evaluates knowledge update methods for legal LLMs across five dimensions. Specifically, we categorize the needs of knowledge updates in the legal domain with the help of legal professionals, and then hire annotators from law schools to create synthetic updates to the Chinese Criminal and Civil Code as well as sets of questions of which the answers would change after the updates. Through a comprehensive evaluation of state-of-the-art knowledge update methods, we reveal a notable gap between existing knowledge update methods and the unique needs of the legal domain, emphasizing the need for further research and development of knowledge update mechanisms tailored for legal LLMs.

📄 PDF Abstract BibTeX arXiv:2407.14192

Code (1)

bebr2/lekube 공식 구현 pytorch

Tasks

Legal Reasoning

Similar Papers 제목 키워드 기반

LexEval: A Comprehensive Chinese Legal Benchmark for Evaluating Large Language Models

2024-09-30 · Haitao Li, You Chen, Qingyao Ai, Yueyue Wu 외

Large language models (LLMs) have made significant progress in natural language processing tasks and demonstrate considerable potential in the legal domain. However, legal applications demand high standards of accuracy, …

Fairness

LawBench: Benchmarking Legal Knowledge of Large Language Models

2023-09-28 · Zhiwei Fei, Xiaoyu Shen, Dawei Zhu, Fengzhe Zhou 외

Large language models (LLMs) have demonstrated strong capabilities in various aspects. However, when applying them to the highly specialized, safe-critical legal domain, it is unclear how much legal knowledge they posses…

ArticlesBenchmarkingMemorizationMulti-Label Classification+1

ArabLegalEval: A Multitask Benchmark for Assessing Arabic Legal Knowledge in Large Language Models

2024-08-15 · Faris Hijazi, Somayah AlHarbi, Abdulaziz AlHussein, Harethah Abu Shairah 외

The rapid advancements in Large Language Models (LLMs) have led to significant improvements in various natural language processing tasks. However, the evaluation of LLMs' legal knowledge, particularly in non-English lang…

In-Context LearningMMLU

Hybrid Retrieval-Augmented Generation Agent for Trustworthy Legal Question Answering in Judicial Forensics

2025-11-03 · Yueqing Xi, Yifan Bai, Huasen Luo, Weiliang Wen 외 arxiv

As artificial intelligence permeates judicial forensics, ensuring the veracity and traceability of legal question answering (QA) has become critical. Conventional large language models (LLMs) are prone to hallucination, …

Question Answering

LawPal : A Retrieval Augmented Generation Based System for Enhanced Legal Accessibility in India

2025-02-23 · Dnyanesh Panchal, Aaryan Gole, Vaibhav Narute, Raunak Joshi

Access to legal knowledge in India is often hindered by a lack of awareness, misinformation and limited accessibility to judicial resources. Many individuals struggle to navigate complex legal frameworks, leading to the …

ChatbotInformation RetrievalMisinformationNavigate+3