paper-with-me

홈 › Papers

Agent Benchmarks Fail Public Sector Requirements

2026-01-28 · Jonathan Rystrøm, Chris Schmitz, Karolina Korgul, Jan Batzner, Chris Russell arxiv

Deploying Large Language Model-based agents (LLM agents) in the public sector requires assuring that they meet the stringent legal, procedural, and structural requirements of public-sector institutions. Practitioners and researchers often turn to benchmarks for such assessments. However, it remains unclear what criteria benchmarks must meet to ensure they adequately reflect public-sector requirements, or how many existing benchmarks do so. In this paper, we first define such criteria based on a first-principles survey of public administration literature: benchmarks must be \emph{process-based}, \emph{realistic}, \emph{public-sector-specific} and report \emph{metrics} that reflect the unique requirements of the public sector. We analyse more than 1,300 benchmark papers for these criteria using an expert-validated LLM-assisted pipeline. Our results show that no single benchmark meets all of the criteria. Our findings provide a call to action for both researchers to develop public sector-relevant benchmarks and for public-sector officials to apply these criteria when evaluating their own agentic use cases.

📄 PDF Abstract BibTeX arXiv:2601.20617

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

From Global Benchmarks to Local Evaluations: Benchmarking LLMs for the German Public Sector

2026-08-18 · Camilla Dalerci, Thilo Michael, Robin Schaefer, Daniel Weinland arxiv

Public institutions face a persistent challenge in selecting LLMs suited to their specific context. Existing benchmarks, however, are of limited use as they primarily reflect English-language and US-centric settings, and…

AI Deployment and Cyber Governance Failures in Public-Sector Organizations: A Typological Analysis

2026-07-28 · Md Salahuddin, James Rooney, Fida Hasan arxiv

The intersection of artificial intelligence adoption, cybersecurity governance, and public sector institutional constraints has not been examined as a unified analytical problem in the existing literature. Studies addres…

Generating and Evaluating Sustainable Procurement Criteria for the Swiss Public Sector using In-Context Prompting with Large Language Models

2026-03-23 · Yingqiang Gao, Veton Matoshi, Luca Rolshoven, Tilia Ellendorff 외 arxiv

Public procurement refers to the process by which public sector institutions, such as governments, municipalities, and publicly funded bodies, acquire goods and services. Swiss law requires the integration of ecological,…

EvoCode-Bench: Evaluating Coding Agents in Multi-Turn Iterative Interactions

2026-05-22 · Haiyang Shen, Xuanzhong Chen, Wendong Xu, Yun Ma 외 arxiv

Coding agents are increasingly used as iterative development partners, but most benchmarks still evaluate one specification followed by one final assessment. This leaves out a basic question: can an agent keep its own co…

Oversight Structures for Agentic AI in Public-Sector Organizations

2025-06-05 · Chris Schmitz, Jonathan Rystrøm, Jan Batzner

This paper finds that the introduction of agentic AI systems intensifies existing challenges to traditional public sector oversight mechanisms -- which rely on siloed compliance units and episodic approvals rather than c…