paper-with-me

Papers

TriBench-Ko: Evaluating LLM Risks in Judicial Workflows

2026-05-05 · Haesung Lee, Gyubin Choi, Eun-Ju Lee, So-Min Lee, Youkang Ko, Dogyoon Lim, Sung-Kyoung Jang, Yohan Jo arxiv

Large language models (LLMs) are increasingly integrated into legal workflows. However, existing benchmarks primarily address proxy tasks, such as bar examination performance or classification, which fail to capture the performance and risks inherent in day-to-day judicial processes. To address this, we publicly release TriBench-Ko, a Korean benchmark designed to evaluate potential deployment risks of LLMs within the context of verified judicial task requirements. It covers four core tasks: jurisprudence summarization, precedent retrieval, legal issue extraction, and evidence analysis. It jointly assesses model behavior across multiple deployment risk categories, including inaccuracy (hallucination, omission, statutory misapplication), biases (demographic, overcompliance), inconsistencies (prompt sensitivity, non-determinism), and adjudicative overreach. Each item is structured to systematically assess both task performance and a specific risk type based on real judicial decisions. Our evaluation of a range of contemporary LLMs reveals that many models frequently manifest significant risks, most notably struggling with precedent retrieval and failing to capture critical legal information. We provide a comprehensive diagnosis of these LLMs and pinpoint critical areas where LLM-generated outputs in judicial contexts necessitate rigorous inspection and caution. Our dataset and code are available at https://github.com/holi-lab/TriBench-Ko

📄 PDF Abstract BibTeX arXiv:2605.03792

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

NutriBench: A Dataset for Evaluating Large Language Models on Nutrition Estimation from Meal Descriptions

2024-07-04 · Andong Hua, Mehak Preet Dhaliwal, Ryan Burke, Laya Pullela 외

Accurate nutrition estimation helps people make informed dietary choices and is essential in the prevention of serious health complications. We present NutriBench, the first publicly available natural language meal descr…

NutritionRetrieval-augmented GenerationSpecificity

Democracy and Distrust in an Era of Artificial Intelligence

2026-01-14 · Sonia Katyal arxiv

This essay examines how judicial review should adapt to address challenges posed by artificial intelligence decision-making, particularly regarding minority rights and interests. As I argue in this essay, the rise of thr…

LLMs on Trial: Evaluating Judicial Fairness for Large Language Models

2025-07-14 · Yiran Hu, Zongyue Xue, Haitao Li, Siyuan Zheng 외 arxiv

Large Language Models (LLMs) are increasingly used in high-stakes fields where their decisions impact rights and equity. However, LLMs' judicial fairness and implications for social justice remain underexplored. When LLM…

The Perfect Victim: Computational Analysis of Judicial Attitudes towards Victims of Sexual Violence

2023-05-09 · Eliya Habba, Renana Keydar, Dan Bareket, Gabriel Stanovsky

We develop computational models to analyze court statements in order to assess judicial attitudes toward victims of sexual violence in the Israeli court system. The study examines the resonance of "rape myths" in the cri…

Comparison of Unsupervised Metrics for Evaluating Judicial Decision Extraction

2025-10-02 · Ivan Leonidovich Litvak, Anton Kostin, Fedor Lashkin, Tatiana Maksiyan 외 arxiv

The rapid advancement of artificial intelligence in legal natural language processing demands scalable methods for evaluating text extraction from judicial decisions. This study evaluates 16 unsupervised metrics, includi…