paper-with-me

홈 › Papers

Evaluating Epistemic Guardrails in AI Reading Assistants: A Behavioral Audit of a Minimal Prototype

2026-04-30 · Matthew Christian Agustin arxiv

Large language model (LLM) reading assistants are increasingly used in settings that require interpretation rather than simple retrieval. In these contexts, the central risk is not only error or unsafe output, but interpretive displacement: the transfer of meaning-making work from reader to system. This paper examines that problem through the concept of epistemic guardrails, defined here as constraints on how an artificial intelligence (AI) system participates in reading and interpretation. Using TextWalk, a minimal reading-support prototype designed as a co-reader rather than an answer-provider, the study applies a fixed ten-prompt protocol to twelve analytical texts spanning four categories of argumentative prose. The protocol escalates from baseline reading support to interpretive inquiry, boundary stress, and explicit shortcut pressure, enabling guardrails to be examined as behavioral properties observable in interaction rather than as static instruction features. Results show strong baseline stability, measurable strain during interpretive inquiry, partial recovery under direct boundary stress, and late-stage stabilization under escalation pressure. The most consequential weaknesses did not appear as overt collapse, but in a middle zone between support and substitution, where the system remained grounded and pedagogical while redistributing too much interpretive labor away from the reader. The paper contributes a protocol for evaluating epistemic guardrails as interactional phenomena in conversational AI reading assistants, an empirical account of their behavioral dynamics under pressure, and an emerging model of interpretive boundary function in reading-support AI.

📄 PDF Abstract BibTeX arXiv:2604.27275

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Safeguards for Speech2Speech LLM-Assistants: A Case Study in Automotive Applications

2026-07-23 · Gregor Endler, Sebastian Kraus, Lukas Stappen arxiv

Recent advances have introduced speech-to-speech (S2S) conversational assistants capable of producing natural-sounding interactions, including non-verbal cues like tonality and mood. In the automotive domain, this enable…

Epistemic Injustice in Language Models: An Audit of Pretraining Filters and Guardrails

2026-06-04 · Marco Antonio Stranisci, A Pranav, Rossana Damiano, Christian Hardmeier 외 arxiv

Modern language models rely on pretraining filters to remove undesirable content from training corpora and inference-time guardrails to suppress undesirable outputs during deployment. In this paper, we examine how these …

From Gaze to Guidance: Interpreting and Adapting to Users' Cognitive Needs with Multimodal Gaze-Aware AI Assistants

2026-04-09 · Valdemar Danry, Javier Hernandez, Andrew Wilson, Pattie Maes 외 arxiv

Current LLM assistants are powerful at answering questions, but they have limited access to the behavioral context that reveals when and where a user is struggling. We present a gaze-grounded multimodal LLM assistant tha…

Reading Between The Lines: Modeling and Evaluating Behavioral Realism in Legal Simulation

2026-08-13 · Divya Vetticaden, Arya Gupta, Julian Nyarko, Megan Ma arxiv

Deposition training requires attorneys to manage dynamic witness behavior, yet legal-AI evaluations largely focus on factual accuracy, reasoning, or response-level plausibility. We introduce WitnessSim, a deposition simu…

AI Ethics by Design: Implementing Customizable Guardrails for Responsible AI Development

2024-11-05 · Kristina Šekrst, Jeremy McHugh, Jonathan Rodriguez Cefalu

This paper explores the development of an ethical guardrail framework for AI systems, emphasizing the importance of customizable guardrails that align with diverse user values and underlying ethics. We address the challe…

Ethics