paper-with-me

홈 › Papers

HalluGraph: Auditable Hallucination Detection for Legal RAG Systems via Knowledge Graph Alignment

2025-12-01 · Valentin Noël, Elimane Yassine Seidou, Charly Ken Capo-Chichi, Ghanem Amari arxiv

Legal AI systems powered by retrieval-augmented generation (RAG) face a critical accountability challenge: when an AI assistant cites case law, statutes, or contractual clauses, practitioners need verifiable guarantees that generated text faithfully represents source documents. Existing hallucination detectors rely on semantic similarity metrics that tolerate entity substitutions, a dangerous failure mode when confusing parties, dates, or legal provisions can have material consequences. We introduce HalluGraph, a graph-theoretic framework that quantifies hallucinations through structural alignment between knowledge graphs extracted from context, query, and response. Our approach produces bounded, interpretable metrics decomposed into \textit{Entity Grounding} (EG), measuring whether entities in the response appear in source documents, and \textit{Relation Preservation} (RP), verifying that asserted relationships are supported by context. On structured control documents, HalluGraph achieves near-perfect discrimination ($>$400 words, $>$20 entities), HalluGraph achieves $AUC = 0.979$, while maintaining robust performance ($AUC \approx 0.89$) on challenging generative legal task, consistently outperforming semantic similarity baselines. The framework provides the transparency and traceability required for high-stakes legal applications, enabling full audit trails from generated assertions back to source passages.

📄 PDF Abstract BibTeX arXiv:2512.01659

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic SimilarityKnowledge Graphs

Similar Papers 제목 키워드 기반

Hybrid Retrieval-Augmented Generation Agent for Trustworthy Legal Question Answering in Judicial Forensics

2025-11-03 · Yueqing Xi, Yifan Bai, Huasen Luo, Weiliang Wen 외 arxiv

As artificial intelligence permeates judicial forensics, ensuring the veracity and traceability of legal question answering (QA) has become critical. Conventional large language models (LLMs) are prone to hallucination, …

Question Answering

How Much Do Legal RAG Systems Still Hallucinate?

2026-08-14 · Souvick Das, Sallam Abualhaija, Domenico Bianculli arxiv

Hallucination is a major challenge for retrieval-augmented generation (RAG) systems in the legal domain, where ungrounded answers can lead to serious consequences. To better understand this problem, we conduct a fine-gra…

Who Checks the Citations? Benchmarking Legal Hallucination Detection

2026-06-19 · Patty Liu, Dominik Stammbach, Peter Henderson arxiv

Attorneys, judges, and pro se filers increasingly use AI to draft legal documents, yet these tools frequently fabricate citations. Despite predictions that newer models would hallucinate less or that court sanctions woul…

Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools

2024-05-30 · Varun Magesh, Faiz Surani, Matthew Dahl, Mirac Suzgun 외

Legal practice has witnessed a sharp rise in products incorporating artificial intelligence (AI). Such tools are designed to assist with a wide range of core legal tasks, from search and summarization of caselaw to docum…

HallucinationRAGRetrieval-augmented Generation

LegalWiz: A Multi-Agent Generation Framework for Contradiction Detection in Legal Documents

2025-10-03 · Ananya Mantravadi, Shivali Dalmia, Olga Pospelova, Abhishek Mukherji 외 arxiv

Retrieval-Augmented Generation (RAG) integrates large language models (LLMs) with external sources, but unresolved contradictions in retrieved evidence often lead to hallucinations and legally unsound outputs. Benchmarks…