paper-with-me

Papers

A Multi-Task Benchmark for Korean Legal Language Understanding and Judgement Prediction

2022-06-10 · Wonseok Hwang, Dongjun Lee, Kyoungyeon Cho, Hanuhl Lee, Minjoon Seo

The recent advances of deep learning have dramatically changed how machine learning, especially in the domain of natural language processing, can be applied to legal domain. However, this shift to the data-driven approaches calls for larger and more diverse datasets, which are nevertheless still small in number, especially in non-English languages. Here we present the first large-scale benchmark of Korean legal AI datasets, LBOX OPEN, that consists of one legal corpus, two classification tasks, two legal judgement prediction (LJP) tasks, and one summarization task. The legal corpus consists of 147k Korean precedents (259M tokens), of which 63k are sentenced in last 4 years and 96k are from the first and the second level courts in which factual issues are reviewed. The two classification tasks are case names (11.3k) and statutes (2.8k) prediction from the factual description of individual cases. The LJP tasks consist of (1) 10.5k criminal examples where the model is asked to predict fine amount, imprisonment with labor, and imprisonment without labor ranges for the given facts, and (2) 4.7k civil examples where the inputs are facts and claim for relief and outputs are the degrees of claim acceptance. The summarization task consists of the Supreme Court precedents and the corresponding summaries (20k). We also release realistic variants of the datasets by extending the domain (1) to infrequent case categories in case name (31k examples) and statute (17.7k) classification tasks, and (2) to long input sequences in the summarization task (51k). Finally, we release LCUBE, the first Korean legal language model trained on the legal corpus from this study. Given the uniqueness of the Law of South Korea and the diversity of the legal tasks covered in this work, we believe that LBOX OPEN contributes to the multilinguality of global legal research. LBOX OPEN and LCUBE will be publicly available.

📄 PDF Abstract BibTeX arXiv:2206.05224

Code (1)

lbox-kr/lbox-open 공식 구현 pytorch

Tasks

Language Modelling

Similar Papers 제목 키워드 기반

Developing a Pragmatic Benchmark for Assessing Korean Legal Language Understanding in Large Language Models

2024-10-11 · Yeeun Kim, Young Rok Choi, Eunkyung Choi, Jinhwan Choi 외

Large language models (LLMs) have demonstrated remarkable performance in the legal domain, with GPT-4 even passing the Uniform Bar Exam in the U.S. However their efficacy remains limited for non-standardized tasks and ta…

Legal ReasoningRAGRetrieval-augmented Generation

CALRK-Bench: Evaluating Context-Aware Legal Reasoning in Korean Law

2026-03-27 · JiHyeok Jung, TaeYoung Yoon, HyunSouk Cho arxiv

Legal reasoning requires not only the application of legal rules but also an understanding of the context in which those rules operate. However, existing legal benchmarks primarily evaluate rule application under the ass…

Legal Reasoning

LegalMidm: Use-Case-Driven Legal Domain Specialization for Korean Large Language Model

2026-04-28 · Youngjoon Jang, Chanhee Park, Hyeonseok Moon, Young-kyoung Ham 외 arxiv

In recent years, the rapid proliferation of open-source large language models (LLMs) has spurred efforts to turn general-purpose models into domain specialists. However, many domain-specialized LLMs are developed using d…

Korean Canonical Legal Benchmark: Toward Knowledge-Independent Evaluation of LLMs' Legal Reasoning Capabilities

2025-12-31 · Hongseok Oh, Wonseok Hwang, Kyoung-Woon On arxiv

We introduce the Korean Canonical Legal Benchmark (KCL), a benchmark designed to assess language models' legal reasoning capabilities independently of domain-specific knowledge. KCL provides question-level supporting pre…

Legal Reasoning

A Korean Legal Judgment Prediction Dataset for Insurance Disputes

2024-01-26 · Alice Saebom Kwak, Cheonkam Jeong, Ji Weon Lim, Byeongcheol Min

This paper introduces a Korean legal judgment prediction (LJP) dataset for insurance disputes. Successful LJP models on insurance disputes can benefit insurance companies and their customers. It can save both sides' time…

Sentence