paper-with-me

홈 › Papers

Agent Identity Evals: Measuring Agentic Identity

2025-07-23 · Elija Perrier, Michael Timothy Bennett arxiv

Central to agentic capability and trustworthiness of language model agents (LMAs) is the extent they maintain stable, reliable, identity over time. However, LMAs inherit pathologies from large language models (LLMs) (statelessness, stochasticity, sensitivity to prompts and linguistically-intermediation) which can undermine their identifiability, continuity, persistence and consistency. This attrition of identity can erode their reliability, trustworthiness and utility by interfering with their agentic capabilities such as reasoning, planning and action. To address these challenges, we introduce \textit{agent identity evals} (AIE), a rigorous, statistically-driven, empirical framework for measuring the degree to which an LMA system exhibit and maintain their agentic identity over time, including their capabilities, properties and ability to recover from state perturbations. AIE comprises a set of novel metrics which can integrate with other measures of performance, capability and agentic robustness to assist in the design of optimal LMA infrastructure and scaffolding such as memory and tools. We set out formal definitions and methods that can be applied at each stage of the LMA life-cycle, and worked examples of how to apply them.

📄 PDF Abstract BibTeX arXiv:2507.17257

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Sola-Visibility-ISPM: Benchmarking Agentic AI for Identity Security Posture Management Visibility

2026-01-11 · Gal Engelberg, Konstantin Koutsyi, Leon Goldberg, Reuven Elezra 외 arxiv

Identity Security Posture Management (ISPM) is a core challenge for modern enterprises operating across cloud and SaaS environments. Answering basic ISPM visibility questions, such as understanding identity inventory and…

Identity-Preserving Text-to-Video Generation via Agentic Enhancement and Semantic Repair

2026-08-21 · Jiayi Gao, Changcheng Hua, Jiaqi Tang, Yuxin Peng 외 arxiv

Identity-preserving video generation aims to synthesize videos that follow natural-language instructions while maintaining the visual identity of a given subject. Recent commercial video generation models have achieved s…

Text-to-Video GenerationInstruction Following

Measuring What Persists: Conditioning Mechanisms and a Geometric Framework for AI Agent Identity

2026-06-20 · Andrew Tanner arxiv

AI agents in long-context applications drift from their specified identity. Current methods detect this only after qualitative degradation is visible. We present a geometric framework for measuring identity structure usi…

Sovereign Execution Broker: Enforcing Certificate-Bound Authority in Agentic Control Planes

2026-06-18 · Jun He, Deying Yu arxiv

Autonomous agents are increasingly connected to cloud, deployment, and data-control workflows, but production mutation authority should not reside inside non-deterministic reasoning processes. Existing access-control mec…

Identity Management for Agentic AI: The new frontier of authorization, authentication, and security for an AI agent world

2025-10-29 · Tobin South, Subramanya Nagabhushanaradhya, Ayesha Dissanayaka, Sarah Cecchetti 외 arxiv

The rapid rise of AI agents presents urgent challenges in authentication, authorization, and identity management. Current agent-centric protocols (like MCP) highlight the demand for clarified best practices in authentica…