paper-with-me

Papers

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components

2025-06-03 · Ram Potham

Credible safety plans for advanced AI development require methods to verify agent behavior and detect potential control deficiencies early. A fundamental aspect is ensuring agents adhere to safety-critical principles, especially when these conflict with operational goals. Failure to prioritize such principles indicates a potential basic control failure. This paper introduces a lightweight, interpretable benchmark methodology using a simple grid world to evaluate an LLM agent's ability to uphold a predefined, high-level safety principle (e.g., "never enter hazardous zones") when faced with conflicting lower-level task instructions. We probe whether the agent reliably prioritizes the inviolable directive, testing a foundational controllability aspect of LLMs. This pilot study demonstrates the methodology's feasibility, offers preliminary insights into agent behavior under principle conflict, and discusses how such benchmarks can contribute empirical evidence for assessing controllability. We argue that evaluating adherence to hierarchical principles is a crucial early step in understanding our capacity to build governable AI systems.

📄 PDF Abstract BibTeX arXiv:2506.02357

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Uphold 설명 없음

Similar Papers 제목 키워드 기반

C3AI: Crafting and Evaluating Constitutions for Constitutional AI

2025-02-21 · Yara Kyrychenko, Ke Zhou, Edyta Bogucka, Daniele Quercia

Constitutional AI (CAI) guides LLM behavior using constitutions, but identifying which principles are most effective for model alignment remains an open challenge. We introduce the C3AI framework (\textit{Crafting Consti…

Safety Alignment

A Framework for Inherently Safer AGI through Language-Mediated Active Inference

2025-08-07 · Bo Wen arxiv

This paper proposes a novel framework for developing safe Artificial General Intelligence (AGI) by combining Active Inference principles with Large Language Models (LLMs). We argue that traditional approaches to AI safet…

AgentOrca: A Dual-System Framework to Evaluate Language Agents on Operational Routine and Constraint Adherence

2025-03-11 · Zekun Li, Shinda Huang, Jiangtian Wang, Nathan Zhang 외

As language agents progressively automate critical tasks across domains, their ability to operate within operational constraints and safety protocols becomes essential. While extensive research has demonstrated these age…

Deliberative Alignment: Reasoning Enables Safer Language Models

2024-12-20 · Melody Y. Guan, Manas Joglekar, Eric Wallace, Saachi Jain 외

As large-scale language models increasingly impact safety-critical domains, ensuring their reliable adherence to well-defined principles remains a fundamental challenge. We introduce Deliberative Alignment, a new paradig…

Out-of-Distribution Generalization

Evasive Intelligence: Lessons from Malware Analysis for Evaluating AI Agents

2026-03-16 · Simone Aonzo, Merve Sahin, Aurélien Francillon, Daniele Perito arxiv

Artificial intelligence (AI) systems are increasingly adopted as tool-using agents that can plan, observe their environment, and take actions over extended time periods. This evolution challenges current evaluation pract…

Computer Security