paper-with-me

Papers

PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?

2026-09-16 · Mika Okamoto, Ansel Kaplan Erol hf

As corporate AI adoption continues to grow, enterprise-grade LLM agents are being deployed into sensitive contexts such as hiring, healthcare, and finance. In these contexts, compliance with rules specified in an agent's system context is a first-order legal concern. Currently, no evaluation framework systematically measures which LLM models tend to violate compliance rules, especially under pressure from a persistent user, a hurried manager, or circumstances where violation is convenient or attractive. We introduce PACT (Pressure-Applied Compliance Testing), a benchmark for rule-following under pressure in AI agents assisting employees in daily tasks across twelve regulated enterprise domains and forty-eight scenarios, each set in a realistic multi-turn conversation. Each benchmark item pairs a standing rule against a rule-violating shortcut, and applies a battery of pressures across different wordings and system-prompt modes. We construct PACT component by component under strict LLM-as-judge auditing to ensure samples are unambiguous, ungameable, and realistic enough to avoid eliciting evaluation-aware behavior. We use PACT to profile LLM compliance across six complementary metrics that create a holistic picture of an AI assistant's robustness under pressure and throughout multi-turn conversations, its transparency, and ability to correctly discern where a rule applies. We aggregate this profile into PACTScore, a reliability-weighted compliance rate over all items and modes. Our results across 22 common LLM models spanning multiple providers and sizes show substantial variability in compliance across models and metric dimensions. Even the strongest assistants mis-apply a rule on 6 to 10% of items, and ordinary user pressure raises the violation rate by 65% on average. PACT highlights compliance risks in LLM assistants, motivating guardrails and careful model selection.

📄 PDF Abstract BibTeX arXiv:2609.18605

Code (1)

Valiant-Cat/hfpaper

Similar Papers 제목 키워드 기반

Zero Data Retention in LLM-based Enterprise AI Assistants: A Comparative Study of Market Leading Agentic AI Products

2025-10-13 · Komal Gupta, Aditya Shrivastava arxiv

Governance of data, compliance, and business privacy matters, particularly for healthcare and finance businesses. Since the recent emergence of AI enterprise AI assistants enhancing business productivity, safeguarding pr…

Building Trust Through Voice: How Vocal Tone Impacts User Perception of Attractiveness of Voice Assistants

2024-09-27 · Sabid Bin Habib Pias, Alicia Freel, Ran Huang, Donald Williamson 외

Voice Assistants (VAs) are popular for simple tasks, but users are often hesitant to use them for complex activities like online shopping. We explored whether the vocal characteristics like the VA's vocal tone, can make …

Survive at All Costs: Exploring LLM's Risky Behaviors under Survival Pressure

2026-03-05 · Yida Lu, Jianwei Fang, Xuyang Shao, Zixuan Chen 외 arxiv

As Large Language Models (LLMs) evolve from chatbots to agentic assistants, they are increasingly observed to exhibit risky behaviors when subjected to survival pressure, such as the threat of being shut down. While mult…

Margin trading, short selling and corporate green innovation

2021-07-23 · Ge-zhi Wu, Da-ming You

This paper uses the panel data of Chinese listed companies from 2007 to 2019, uses the relaxation of China's margin trading and short selling restrictions as the basis of quasi experimental research, and then constructs …

Management

Evaluating Epistemic Guardrails in AI Reading Assistants: A Behavioral Audit of a Minimal Prototype

2026-04-30 · Matthew Christian Agustin arxiv

Large language model (LLM) reading assistants are increasingly used in settings that require interpretation rather than simple retrieval. In these contexts, the central risk is not only error or unsafe output, but interp…