paper-with-me

홈 › Papers

A Benchmark for Gap and Overlap Analysis as a Test of KG Task Readiness

2026-04-12 · Maruf Ahmed Mridul, Rohit Kapa, Oshani Seneviratne arxiv

Task-oriented evaluation of knowledge graph (KG) quality increasingly asks whether an ontology-based representation can answer the competency questions that users actually care about, in a manner that is reproducible, explainable, and traceable to evidence. This paper adopts that perspective and focuses on gap and overlap analysis for policy-like documents (e.g., insurance contracts), where given a scenario, which documents support it (overlap) and which do not (gap), with defensible justifications. The resulting gap/overlap determinations are typically driven by genuine differences in coverage and restrictions rather than missing data, making the task a direct test of KG task readiness rather than a test of missing facts or query expressiveness. We present an executable and auditable benchmark that aligns natural-language contract text with a formal ontology and evidence-linked ground truth, enabling systematic comparison of methods. The benchmark includes: (i) ten simplified yet diverse life-insurance contracts reviewed by a domain expert, (ii) a domain ontology (TBox) with an instantiated knowledge base (ABox) populated from contract facts, and (iii) 58 structured scenarios paired with SPARQL queries with contract-level outcomes and clause-level excerpts that justify each label. Using this resource, we compare a text-only LLM baseline that infers outcomes directly from contract text against an ontology-driven pipeline that answers the same scenarios over the instantiated KG, demonstrating that explicit modeling improves consistency and diagnosis for gap/overlap analyses. Although demonstrated for gap and overlap analysis, the benchmark is intended as a reusable template for evaluating KG quality and supporting downstream work such as ontology learning, KG population, and evidence-grounded question answering.

📄 PDF Abstract BibTeX arXiv:2604.10853

Code (0)

등록된 구현이 없습니다.

Tasks

Question Answering

Similar Papers 제목 키워드 기반

Exploratory Visual Analysis for Increasing Data Readiness in Artificial Intelligence Projects

2024-09-05 · Mattias Tiger, Daniel Jakobsson, Anders Ynnerman, Fredrik Heintz 외

We present experiences and lessons learned from increasing data readiness of heterogeneous data for artificial intelligence projects using visual analysis methods. Increasing the data readiness level involves understandi…

REBAR: Reference Ethical Benchmark for Autonomy Readiness

2026-05-18 · Jonathan Diller, David Barnes, Rebekah Bogdanoff, Rhett Collier 외 arxiv

As autonomous systems grow more advanced, objective metrics to evaluate their ethical and legal compliance are critical for informing end users of their limitations and ensuring accountability of those who misuse them. C…

Red Teaming

RAIL: An Automatic Classifier of the Artificial Intelligence Readiness Level

2026-08-13 · Juan Irving Vasquez, Juan Terven, Laura-Ivoone Garay-Jimenez arxiv

Assessing the maturity of artificial intelligence technologies is essential for investment decisions, project management, and policy monitoring, yet the available readiness frameworks are heterogeneous and difficult to a…

Data Readiness Report

2020-10-14 · Shazia Afzal, Rajmohan C, Manish Kesarwani, Sameep Mehta 외

Data exploration and quality analysis is an important yet tedious process in the AI pipeline. Current practices of data cleaning and data readiness assessment for machine learning tasks are mostly conducted in an arbitra…

AutoMLManagementNutrition

The Illusion of Readiness in Health AI

2025-09-22 · Yu Gu, Jingjing Fu, Xiaodong Liu, Jeya Maria Jose Valanarasu 외 arxiv

Large language models have demonstrated remarkable performance in a wide range of medical benchmarks. Yet underneath the seemingly promising results lie salient growth areas, especially in cutting-edge frontiers such as …

Multimodal Reasoning