paper-with-me

홈 › Papers

From Legal Text to Executable Decision Models: Evaluating Structured Representations for Legal Decision Model Generation

2026-04-18 · David Graus arxiv

Transforming legal text into executable decision logic is a longstanding challenge in legal informatics. With the rise of LLMs, this task has gained renewed interest, but remains challenging due to requiring extensive manual coding and evaluation. We use a unique real-world dataset that pairs production-grade decision models with legal text from the Dutch Environment and Planning Act. These models power the Omgevingsloket government platform, where citizens check permit requirements for environmental activities. We study whether intermediate structured representations can improve LLM-based generation of executable decision models from legal text. We compare four input conditions: raw legal text, text enriched with semantic role labels, text enriched with input and output constraints, and text enriched with both. We evaluate along two dimensions: structural evaluation, through similarity to gold decision models with graph kernels and graphs' descriptive statistics, and outcome evaluation, through functional equivalence by executing models on pre-configured test scenarios. Our findings show that I/O constraints provide the dominant improvement (+37-54% similarity over baseline), while semantic role labels show modest improvements. Outcome evaluation shows that generated models match the gold standard on 51-53% of test scenarios, even though generated models are typically smaller and simpler. We find LLMs eliminate redundant pass-through logic that comprises up to 45-55% of nodes. Importantly, structural similarity and outcome equivalence are complementary: structural similarity does not guarantee outcome equivalence, and vice versa. To facilitate reproducibility, we publicly release our dataset of 95 production decision models with associated legal text and all experimental code.

📄 PDF Abstract BibTeX arXiv:2604.17153

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

When Fairness Isn't Statistical: The Limits of Machine Learning in Evaluating Legal Reasoning

2025-06-04 · Claire Barale, Michael Rovatsos, Nehal Bhuta

Legal decisions are increasingly evaluated for fairness, consistency, and bias using machine learning (ML) techniques. In high-stakes domains like refugee adjudication, such methods are often applied to detect disparitie…

ClusteringFairnessLegal Reasoning

An LLM Agentic Approach for Legal-Critical Software: A Case Study for Tax Prep Software

2025-09-16 · Sina Gogani-Khiabani, Ashutosh Trivedi, Diptikalyan Saha, Saeid Tizpaz-Niari arxiv

Large language models (LLMs) show promise for translating natural-language statutes into executable logic, but reliability in legally critical settings remains challenging due to ambiguity and hallucinations. We present …

ARMOR 2025: A Military-Aligned Benchmark for Evaluating Large Language Model Safety Beyond Civilian Contexts

2026-04-30 · Sydney Johns, Heng Jin, Chaoyu Zhang, Y. Thomas Hou 외 arxiv

Large language models (LLMs) are now being explored for defense applications that require reliable and legally compliant decision support. They also hold significant potential to enhance decision making, coordination, an…

Decision Making

Comparison of Unsupervised Metrics for Evaluating Judicial Decision Extraction

2025-10-02 · Ivan Leonidovich Litvak, Anton Kostin, Fedor Lashkin, Tatiana Maksiyan 외 arxiv

The rapid advancement of artificial intelligence in legal natural language processing demands scalable methods for evaluating text extraction from judicial decisions. This study evaluates 16 unsupervised metrics, includi…

LLM4SFC: Sequential Function Chart Generation via Large Language Models

2025-12-07 · Ofek Glick, Vladimir Tchuiev, Marah Ghoummaid, Michal Moshkovitz 외 arxiv

While Large Language Models (LLMs) are increasingly used for synthesizing textual PLC programming languages like Structured Text (ST) code, other IEC 61131-3 standard graphical languages like Sequential Function Charts (…