paper-with-me

Papers

Beyond single-channel agentic benchmarking

2026-02-05 · Nelu D. Radpour arxiv

Contemporary benchmarks for agentic artificial intelligence (AI) frequently evaluate safety through isolated task-level accuracy thresholds, implicitly treating autonomous systems as single points of failure. This single-channel paradigm diverges from established principles in safety-critical engineering, where risk mitigation is achieved through redundancy, diversity of error modes, and joint system reliability. This paper argues that evaluating AI agents in isolation systematically mischaracterizes their operational safety when deployed within human-in-the-loop environments. Using a recent laboratory safety benchmark as a case study demonstrates that even imperfect AI systems can nonetheless provide substantial safety utility by functioning as redundant audit layers against well-documented sources of human failure, including vigilance decrement, inattentional blindness, and normalization of deviance. This perspective reframes agentic safety evaluation around the reliability of the human-AI dyad rather than absolute agent accuracy, with a particular emphasis on uncorrelated error modes as the primary determinant of risk reduction. Such a shift aligns AI benchmarking with established practices in other safety-critical domains and offers a path toward more ecologically valid safety assessments.

📄 PDF Abstract BibTeX arXiv:2602.18456

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Beyond Black-Box Benchmarking: Observability, Analytics, and Optimization of Agentic Systems

2025-03-09 · Dany Moshkovich, Hadar Mulian, Sergey Zeltyn, Natti Eder 외

The rise of agentic AI systems, where agents collaborate to perform diverse tasks, poses new challenges with observing, analyzing and optimizing their behavior. Traditional evaluation and benchmarking approaches struggle…

Benchmarking

Do We Still Need GraphRAG? Benchmarking RAG and GraphRAG for Agentic Search Systems

2026-04-01 · Dongzhe Fan, Zheyi Xue, Siyuan Liu, Qiaoyu Tan arxiv

Retrieval-augmented generation (RAG) and its graph-based extensions (GraphRAG) are effective paradigms for improving large language model (LLM) reasoning by grounding generation in external knowledge. However, most exist…

Question Answering

Problem Reductions at Scale: Agentic Integration of Computationally Hard Problems

2026-04-13 · Xi-Wei Pan, Shi-Wen An, Jin-Guo Liu arxiv

Solving an NP-hard optimization problem often requires reformulating it for a specific solver -- quantum hardware, a commercial optimizer, or a domain heuristic. A tool for polynomial-time reductions between hard problem…

Agentic AI Systems in Electrical Power Systems Engineering: Current State-of-the-Art and Challenges

2025-11-18 · Soham Ghosh, Gaurav Mittal arxiv

Agentic AI systems have recently emerged as a critical and transformative approach in artificial intelligence, offering capabilities that extend far beyond traditional AI agents and contemporary generative AI models. Thi…

ATOD: An Evaluation Framework and Benchmark for Agentic Task-Oriented Dialogue Systems

2026-01-17 · Yifei Zhang, Hooshang Nayyeri, Rinat Khaziev, Emine Yilmaz 외 arxiv

Recent advances in task-oriented dialogue (TOD) systems, driven by large language models (LLMs) with extensive API and tool integration, have enabled conversational agents to coordinate interleaved goals, maintain long-h…

Task-Oriented Dialogue SystemsDialogue Generation