paper-with-me

Papers

Transparent, Evaluable, and Accessible Data Agents: A Proof-of-Concept Framework

2025-09-28 · Nooshin Bahador arxiv

This article presents a modular, component-based architecture for developing and evaluating AI agents that bridge the gap between natural language interfaces and complex enterprise data warehouses. The system directly addresses core challenges in data accessibility by enabling non-technical users to interact with complex data warehouses through a conversational interface, translating ambiguous user intent into precise, executable database queries to overcome semantic gaps. A cornerstone of the design is its commitment to transparent decision-making, achieved through a multi-layered reasoning framework that explains the "why" behind every decision, allowing for full interpretability by tracing conclusions through specific, activated business rules and data points. The architecture integrates a robust quality assurance mechanism via an automated evaluation framework that serves multiple functions: it enables performance benchmarking by objectively measuring agent performance against golden standards, and it ensures system reliability by automating the detection of performance regressions during updates. The agent's analytical depth is enhanced by a statistical context module, which quantifies deviations from normative behavior, ensuring all conclusions are supported by quantitative evidence including concrete data, percentages, and statistical comparisons. We demonstrate the efficacy of this integrated agent-development-with-evaluation framework through a case study on an insurance claims processing system. The agent, built on a modular architecture, leverages the BigQuery ecosystem to perform secure data retrieval, apply domain-specific business rules, and generate human-auditable justifications. The results confirm that this approach creates a robust, evaluable, and trustworthy system for deploying LLM-powered agents in data-sensitive, high-stakes domains.

📄 PDF Abstract BibTeX arXiv:2509.24127

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Argumentation Semantics for Prioritised Default Logic

2015-06-26 · Anthony P. Young, Sanjay Modgil, Odinaldo Rodrigues

We endow prioritised default logic (PDL) with argumentation semantics using the ASPIC+ framework for structured argumentation, and prove that the conclusions of the justified arguments are exactly the prioritised default…

Federated Learning using Smart Contracts on Blockchains, based on Reward Driven Approach

2021-07-19 · Monik Raj Behera, Sudhir Upadhyay, Suresh Shetty

Over the recent years, Federated machine learning continues to gain interest and momentum where there is a need to draw insights from data while preserving the data provider's privacy. However, one among other existing c…

Federated Learning

HierSVA: A Data Synthesis Pipeline, Dataset, and Benchmark for LLM-Driven Hierarchical Hardware Formal Verification

2026-06-09 · Maohua Nie, Jiang Zhu, Jingqun Zhang, Zhichen Zeng 외 arxiv

We present HierSVA, an integrated suite that combines a pipeline, dataset, and benchmark for LLM-driven hierarchical hardware formal verification. HierSVA-SP pairs an RTL preprocessing toolchain with an LLM-in-the-loop f…

Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment

2026-08-24 · Stephen Chung, Wenyu Du, William J. Wesley hf

We study autonomous mathematical discovery in the Station, an open-world multi-agent environment in which AI agents from different model families pursue a shared research goal without a central coordinator or scripted pi…

Veracity: An Open-Source AI Fact-Checking System

2025-06-18 · Taylor Lynn Curtis, Maximilian Puelma Touzel, William Garneau, Manon Gruaz 외

The proliferation of misinformation poses a significant threat to society, exacerbated by the capabilities of generative AI. This demo paper introduces Veracity, an open-source AI system designed to empower individuals t…

Fact CheckingMisinformationRetrieval