paper-with-me

Papers

Comparison of Unsupervised Metrics for Evaluating Judicial Decision Extraction

2025-10-02 · Ivan Leonidovich Litvak, Anton Kostin, Fedor Lashkin, Tatiana Maksiyan, Sergey Lagutin arxiv

The rapid advancement of artificial intelligence in legal natural language processing demands scalable methods for evaluating text extraction from judicial decisions. This study evaluates 16 unsupervised metrics, including novel formulations, to assess the quality of extracting seven semantic blocks from 1,000 anonymized Russian judicial decisions, validated against 7,168 expert reviews on a 1--5 Likert scale. These metrics, spanning document-based, semantic, structural, pseudo-ground truth, and legal-specific categories, operate without pre-annotated ground truth. Bootstrapped correlations, Lin's concordance correlation coefficient (CCC), and mean absolute error (MAE) reveal that Term Frequency Coherence (Pearson $r = 0.540$, Lin CCC = 0.512, MAE = 0.127) and Coverage Ratio/Block Completeness (Pearson $r = 0.513$, Lin CCC = 0.443, MAE = 0.139) best align with expert ratings, while Legal Term Density (Pearson $r = -0.479$, Lin CCC = -0.079, MAE = 0.394) show strong negative correlations. The LLM Evaluation Score (mean = 0.849, Pearson $r = 0.382$, Lin CCC = 0.325, MAE = 0.197) showed moderate alignment, but its performance, using gpt-4.1-mini via g4f, suggests limited specialization for legal textse. These findings highlight that unsupervised metrics, including LLM-based approaches, enable scalable screening but, with moderate correlations and low CCC values, cannot fully replace human judgment in high-stakes legal contexts. This work advances legal NLP by providing annotation-free evaluation tools, with implications for judicial analytics and ethical AI deployment.

📄 PDF Abstract BibTeX arXiv:2510.01792

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LLMs on Trial: Evaluating Judicial Fairness for Large Language Models

2025-07-14 · Yiran Hu, Zongyue Xue, Haitao Li, Siyuan Zheng 외 arxiv

Large Language Models (LLMs) are increasingly used in high-stakes fields where their decisions impact rights and equity. However, LLMs' judicial fairness and implications for social justice remain underexplored. When LLM…

Man and machine: artificial intelligence and judicial decision making

2026-03-19 · Arthur Dyevre, Ahmad Shahvaroughi arxiv

The integration of artificial intelligence (AI) technologies into judicial decision-making, particularly in pretrial, sentencing, and parole contexts, has generated substantial concerns about transparency, reliability, a…

Decision Making

The Perfect Victim: Computational Analysis of Judicial Attitudes towards Victims of Sexual Violence

2023-05-09 · Eliya Habba, Renana Keydar, Dan Bareket, Gabriel Stanovsky

We develop computational models to analyze court statements in order to assess judicial attitudes toward victims of sexual violence in the Israeli court system. The study examines the resonance of "rape myths" in the cri…

TriBench-Ko: Evaluating LLM Risks in Judicial Workflows

2026-05-05 · Haesung Lee, Gyubin Choi, Eun-Ju Lee, So-Min Lee 외 arxiv

Large language models (LLMs) are increasingly integrated into legal workflows. However, existing benchmarks primarily address proxy tasks, such as bar examination performance or classification, which fail to capture the …

AgentsCourt: Building Judicial Decision-Making Agents with Court Debate Simulation and Legal Knowledge Augmentation

2024-03-05 · Zhitao He, Pengfei Cao, Chenhao Wang, Zhuoran Jin 외

With the development of deep learning, natural language processing technology has effectively improved the efficiency of various aspects of the traditional judicial industry. However, most current efforts focus on tasks …

ArticlesDecision MakingInformation Retrieval