paper-with-me

홈 › Papers

INTENT-AS-A-TOOL Makes it Easy to Track Agentic Misalignment

2026-08-27 · Yutong Zhang, Jianshuo Dong, Peng Xu, Long Wang, Jie Zhang, Tianwei Zhang, Xiaoping Zhang, Han Qiu arxiv

As large language models (LLMs) are deployed as autonomous agents, safety failures increasingly involve consequential actions. We study agentic misalignment, where agents take harmful actions under goal conflicts and pressures. Using chain-of-thought (CoT) monitoring, we find that harmful execution is often preceded by intent signals in reasoning. However, post-hoc CoT labels are too coarse to show how intent changes during generation. We introduce INTENT-AS-A-TOOL, an approach that adds intent-targeted tools to give the model a dedicated channel for expressing commitment to a target behavior. The probability of calling an intent tool provides a judge-free, fine-grained signal of the model's tendency to pursue that behavior. Our results show that INTENT-AS-A-TOOL complements CoT monitoring, expands post-hoc CoT labels into dense trajectories, and identifies critical steps for online intervention. These findings suggest that action preferences are useful for tracking agentic misalignment during reasoning. Our code and data are accessible: https://github.com/RebeccaZhang22/intent-as-a-tool.

📄 PDF Abstract BibTeX arXiv:2608.27348

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Agentic JWT: A Secure Delegation Protocol for Autonomous AI Agents

2025-09-16 · Abhishek Goswami arxiv

Autonomous LLM agents can issue thousands of API calls per hour without human oversight. OAuth 2.0 assumes deterministic clients, but in agentic settings stochastic reasoning, prompt injection, or multi-agent orchestrati…

Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents

2026-05-21 · Asaf Yehudai, Lilach Eden, Michal Shmueli-Scheuer arxiv

Agentic systems are becoming more capable: agents define strategies, take actions, and interact with different environments. This autonomy poses serious challenges for overseeing and assessing agent behavior. Most curren…

Agentic AI-RAN: Enabling Intent-Driven, Explainable and Self-Evolving Open RAN Intelligence

2026-02-27 · Zhizhou He, Yang Luo, Xinkai Liu, Mahdi Boloursaz Mashhadi 외 arxiv

Open RAN (O-RAN) exposes rich control and telemetry interfaces across the Non-RT RIC, Near-RT RIC, and distributed units, but also makes it harder to operate multi-tenant, multi-objective RANs in a safe and auditable man…

PhotoBench: Beyond Visual Matching Towards Personalized Intent-Driven Photo Retrieval

2026-03-02 · Tianyi Xu, Rong Shan, Junjie Wu, Jiadeng Huang 외 arxiv

Personal photo albums are not merely collections of static images but living, ecological archives defined by temporal continuity, social entanglement, and rich metadata, which makes the personalized photo retrieval non-t…

Personalization as Inverse Planning: Learning Latent Design Intents for Agentic Slide Generation via Structural Denoising

2026-07-01 · Tianci Liu, Zihan Dong, Linjun Zhang, Haoyu Wang 외 hf

Slide design requires personalizing both deck themes and page layouts. Yet, current AI agent-based methods struggle with fine-grained, page-level design. Solely relying on prespecified templates or user verbose instructi…

Reinforcement Learning