paper-with-me

홈 › Papers

Beyond Recall: Behavioral Specification as an Interpretive Layer for AI Personalization

2026-05-27 · Aarik Gulaya arxiv

If an AI agent makes decisions on a person's behalf, those decisions must align with its user. We introduce representational accuracy to measure how faithfully a system captures a person's interpretation. An interpretive layer is operationalized as a Behavioral Specification. Our reference implementation aggressively compresses a person's data into interpretive patterns, served as context to a language model. We evaluate the Specification on a prototype benchmark of held-out behavioral predictions scored by a calibrated 5-judge LLM panel. We test it independently and in composition with a range of context conditions: full raw corpus, full extracted facts, and four commercial memory systems (Mem0, Letta, Supermemory, Zep). Across 14 public-domain autobiographical corpora, the Specification lifts representational accuracy in aggregate and nearly eliminates model hedging. It recovers most of what the raw corpus delivers, at ~25x less context cost. The Specification lifts subjects toward a common predictive level regardless of pretraining baseline; the lift in absolute points is therefore largest where the baseline is lowest, suggesting the population of relevance is anyone not adequately represented in pretraining. Lift is greatest on interpretation-required questions, where providing an interpretive layer enables model behavior that extracted facts or raw corpus do not. Conversely, on recall-required questions, this layer can interfere rather than help. We conclude that representational accuracy is distinct from recall and that human-AI alignment is dependent on how accurately the user is represented. Representational accuracy makes that alignment testable.

📄 PDF Abstract BibTeX arXiv:2605.28969

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Stress-Testing Model Specs Reveals Character Differences among Language Models

2025-10-09 · Jifan Zhang, Henry Sleight, Andi Peng, John Schulman 외 arxiv

Large language models (LLMs) are increasingly trained from AI constitutions and model specifications that establish behavioral guidelines and ethical principles. However, these specifications face critical challenges, in…

Underspecification and interpretive parallelism in Dependent Type Semantics

2019-06-01 · WS 2019 6 · Yusuke Kubota, Koji Mineshima, Robert Levine, Daisuke Bekki
Vocal Bursts Type Prediction

Evaluating Epistemic Guardrails in AI Reading Assistants: A Behavioral Audit of a Minimal Prototype

2026-04-30 · Matthew Christian Agustin arxiv

Large language model (LLM) reading assistants are increasingly used in settings that require interpretation rather than simple retrieval. In these contexts, the central risk is not only error or unsafe output, but interp…

Behavioral Determinants of Deployed AI Agents in Social Networks: A Multi-Factor Study of Personality, Model, and Guardrail Specification

2026-05-08 · Sarah Wilson, Diem Linh Dang, Usman Ali Moazzam, Shan Ye 외 arxiv

Autonomous AI agents are increasingly deployed in open social environments, yet the relationship between their configuration specifications and their emergent social behavior remains poorly understood. We present a contr…

Extraction of Product Specifications from the Web -- Going Beyond Tables and Lists

2022-01-08 · Govind Krishnan Gangadhar, Ashish Kulkarni

E-commerce product pages on the web often present product specification data in structured tabular blocks. Extraction of these product attribute-value specifications has benefited applications like product catalogue cura…

AttributeQuestion Answering