paper-with-me

홈 › Papers

COMPASS: A Framework for Evaluating Organization-Specific Policy Alignment in LLMs

2026-01-05 · Dasol Choi, DongGeon Lee, Brigitta Jesica Kartono, Helena Berndt, Taeyoun Kwon, Joonwon Jang, Haon Park, Hwanjo Yu, Minsuk Kahng arxiv

As large language models are deployed in high-stakes enterprise applications, from healthcare to finance, ensuring adherence to organization-specific policies has become essential. Yet existing safety evaluations focus exclusively on universal harms. We present COMPASS (Company/Organization Policy Alignment Assessment), the first systematic framework for evaluating whether LLMs comply with organizational allowlist and denylist policies. We apply COMPASS to eight diverse industry scenarios, generating and validating 5,920 queries that test both routine compliance and adversarial robustness through strategically designed edge cases. Evaluating seven state-of-the-art models, we uncover a fundamental asymmetry: models reliably handle legitimate requests (>95% accuracy) but catastrophically fail at enforcing prohibitions, refusing only 13-40% of adversarial denylist violations. These results demonstrate that current LLMs lack the robustness required for policy-critical deployments, establishing COMPASS as an essential evaluation framework for organizational AI safety.

📄 PDF Abstract BibTeX arXiv:2601.01836

Code (0)

등록된 구현이 없습니다.

Tasks

Adversarial Robustness

Similar Papers 제목 키워드 기반

PolicyGuard: From Organizational Policies to Neuro-SymbolicCompliance Review Engines

2026-06-30 · Sameer Malik, Ayush Singh, Amar Prakash Azad arxiv

Policy-grounded document review requires determining whether a target document complies with organization-specific policies, guidelines, or playbooks. While large language models can assist with policy interpretation and…

A Logic of Agent Organizations

2018-04-28 · Virginia Dignum, Frank Dignum

Organization concepts and models are increasingly being adopted for the design and specification of multi-agent systems. Agent organizations can be seen as mechanisms of social order, created to achieve global (or organi…

A Semantic Framework for Enabling Radio Spectrum Policy Management and Evaluation

2020-11-08 · H. Santos, A. Mulvehill, J. S. Erickson, J. P. McCusker 외

Because radio spectrum is a finite resource, its usage and sharing is regulated by government agencies. These agencies define policies to manage spectrum allocation and assignment across multiple organizations, systems, …

Management

PrivComp-KG : Leveraging Knowledge Graph and Large Language Models for Privacy Policy Compliance Verification

2024-04-30 · Leon Garza, Lavanya Elluri, Anantaa Kotal, Aritran Piplai 외

Data protection and privacy is becoming increasingly crucial in the digital era. Numerous companies depend on third-party vendors and service providers to carry out critical functions within their operations, encompassin…

Language ModellingLarge Language ModelRetrieval-augmented Generation

Abstract Reward Processes: Leveraging State Abstraction for Consistent Off-Policy Evaluation

2024-10-03 · Shreyas Chaudhari, Ameet Deshpande, Bruno Castro da Silva, Philip S. Thomas

Evaluating policies using off-policy data is crucial for applying reinforcement learning to real-world problems such as healthcare and autonomous driving. Previous methods for off-policy evaluation (OPE) generally suffer…

Autonomous DrivingOff-policy evaluation