paper-with-me

Papers

AEMA: Verifiable Evaluation Framework for Trustworthy and Controlled Agentic LLM Systems

2026-01-17 · YenTing Lee, Keerthi Koneru, Zahra Moslemi, Sheethal Kumar, Ramesh Radhakrishnan arxiv

Evaluating large language model (LLM)-based multi-agent systems remains a critical challenge, as these systems must exhibit reliable coordination, transparent decision-making, and verifiable performance across evolving tasks. Existing evaluation approaches often limit themselves to single-response scoring or narrow benchmarks, which lack stability, extensibility, and automation when deployed in enterprise settings at multi-agent scale. We present AEMA (Adaptive Evaluation Multi-Agent), a process-aware and auditable framework that plans, executes, and aggregates multi-step evaluations across heterogeneous agentic workflows under human oversight. Compared to a single LLM-as-a-Judge, AEMA achieves greater stability, human alignment, and traceable records that support accountable automation. Our results on enterprise-style agent workflows simulated using realistic business scenarios demonstrate that AEMA provides a transparent and reproducible pathway toward responsible evaluation of LLM-based multi-agent systems. Keywords Agentic AI, Multi-Agent Systems, Trustworthy AI, Verifiable Evaluation, Human Oversight

📄 PDF Abstract BibTeX arXiv:2601.11903

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A feature-stable and explainable machine learning framework for trustworthy decision-making under incomplete clinical data

2026-02-19 · Justyna Andrys-Olek, Paulina Tworek, Luca Gherardini, Mark W. Ruddock 외 arxiv

Machine learning models are increasingly applied to biomedical data, yet their adoption in high stakes domains remains limited by poor robustness, limited interpretability, and instability of learned features under reali…

Homogenisation of nonlinear blood flow in periodic networks: the limit of small haematocrit heterogeneity

2024-01-16 · Y. Ben-Ami, B. D. Wood, J. M. Pitt-Francis, P. K. Maini 외

In this work we develop a homogenisation methodology to upscale mathematical descriptions of microcirculatory blood flow from the microscale (where individual vessels are resolved) to the macroscopic (or tissue) scale. D…

Closing the AI Trust Gap: The Case for Independent Certification for Trustworthy AI

2026-07-17 · Trisevgeni Papakonstantinou, Cansu Canca, Farah Nanji, Waheedullah Pardess 외 arxiv

Over the past decade, responsible AI (RAI) has produced a substantial body of practice for identifying and mitigating the risks AI poses in high-stakes settings. Yet this work has not produced a market that rewards trust…

Synergistic Perception-Reasoning Governance: Grounding Medical MLLMs with Verifiable Anatomical Evidence

2026-06-30 · Rui Hao, Qiankun Li, Junyuan Mao, Linghao Meng 외 arxiv

Multimodal large language models (MLLMs) show strong promise for clinical VQA and radiology report generation, yet inference-time hallucinations still undermine trustworthy use: models can produce fluent conclusions that…

AI Bill of Materials and Beyond: Systematizing Security Assurance through the AI Risk Scanning (AIRS) Framework

2025-11-16 · Samuel Nathanson, Alexander Lee, Catherine Chen Kieffer, Jared Junkin 외 arxiv

Assurance for artificial intelligence (AI) systems remains fragmented across software supply-chain security, adversarial machine learning, and governance documentation. Existing transparency mechanisms - including Model …