paper-with-me

Papers

Validity Is What You Need

2025-10-31 · Sebastian Benthall, Andrew Clark arxiv

While AI agents have long been discussed and studied in computer science, today's Agentic AI systems are something new. We consider other definitions of Agentic AI and propose a new realist definition. Agentic AI is a software delivery mechanism, comparable to software as a service (SaaS), which puts an application to work autonomously in a complex enterprise setting. Recent advances in large language models (LLMs) as foundation models have driven excitement in Agentic AI. We note, however, that Agentic AI systems are primarily applications, not foundations, and so their success depends on validation by end users and principal stakeholders. The tools and techniques needed by the principal users to validate their applications are quite different from the tools and techniques used to evaluate foundation models. Ironically, with good validation measures in place, in many cases the foundation models can be replaced with much simpler, faster, and more interpretable models that handle core logic. When it comes to Agentic AI, validity is what you need. LLMs are one option that might achieve it.

📄 PDF Abstract BibTeX arXiv:2510.27628

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Falsifying Discriminant Validity of Predictive Algorithms

2026-01-23 · Amanda Coston arxiv

Empirical investigations into unintended model behavior often show that the algorithm is predicting another outcome than what was intended. These exposés highlight the need to identify when algorithms predict unintended …

Causal Inference

Exploring the Use of ChatGPT for a Systematic Literature Review: a Design-Based Research

2024-09-25 · Qian Huang, Qiyun Wang

ChatGPT has been used in several educational contexts,including learning, teaching and research. It also has potential to conduct the systematic literature review (SLR). However, there are limited empirical studies on ho…

Systematic Literature Review

Measuring what Matters: Construct Validity in Large Language Model Benchmarks

2025-11-03 · Andrew M. Bean, Ryan Othniel Kearns, Angelika Romanou, Franziska Sofia Hafner 외 arxiv

Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstract and complex phenomena such as 'safety'…

BenchBrowser: Retrieving Evidence for Evaluating Benchmark Validity

2026-02-25 · Harshita Diddee, Gregory Yauney, Swabha Swayamdipta, Daphne Ippolito arxiv

Do language model benchmarks actually measure what practitioners intend them to ? High-level metadata is too coarse to convey the granular reality of benchmarks: a "poetry" benchmark may never test for haikus, while "ins…

Quantifying the Internal Validity of Weighted Estimands

2024-04-22 · Alexandre Poirier, Tymon Słoczyński

In this paper we study a class of weighted estimands, which we define as parameters that can be expressed as weighted averages of the underlying heterogeneous treatment effects. The popular ordinary least squares (OLS), …

Diagnostic