paper-with-me

홈 › Papers

SafePro: Evaluating the Safety of Professional-Level AI Agents

2026-01-10 · Kaiwen Zhou, Shreedhar Jangam, Ashwin Nagarajan, Tejas Polu, Suhas Oruganti, Chengzhi Liu, Ching-Chen Kuo, Yuting Zheng, Sravana Narayanaraju, Xin Eric Wang arxiv

Large language model-based agents are rapidly evolving from simple conversational assistants into autonomous systems capable of performing complex, professional-level tasks in various domains. While these advancements promise significant productivity gains, they also introduce critical safety risks that remain under-explored. Existing safety evaluations primarily focus on simple, daily assistance tasks, failing to capture the intricate decision-making processes and potential consequences of misaligned behaviors in professional settings. To address this gap, we introduce \textbf{SafePro}, a comprehensive benchmark designed to evaluate the safety alignment of AI agents performing professional activities. SafePro features a dataset of high-complexity tasks across diverse professional domains with safety risks, developed through a rigorous iterative creation and review process. Our evaluation of state-of-the-art AI models reveals significant safety vulnerabilities and uncovers new unsafe behaviors in professional contexts. We further show that these models exhibit both insufficient safety judgment and weak safety alignment when executing complex professional tasks. In addition, we investigate safety mitigation strategies for improving agent safety in these scenarios and observe encouraging improvements. Together, our findings highlight the urgent need for robust safety mechanisms tailored to the next generation of professional AI agents.

📄 PDF Abstract BibTeX arXiv:2601.06663

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SafeProtein: Red-Teaming Framework and Benchmark for Protein Foundation Models

2025-09-03 · Jigang Fan, Zhenghong Zhou, Ruofan Jin, Le Cong 외 arxiv

Proteins play crucial roles in almost all biological processes. The advancement of deep learning has greatly accelerated the development of protein foundation models, leading to significant successes in protein understan…

Prompt Engineering

MBABench: Evaluating LLM Agents on End-to-End Spreadsheet Tasks in Finance

2026-05-21 · Thomson Yen, Julian Poeltl, Harshith Srinivas Gear, Yilin Meng 외 arxiv

LLM agents are increasingly expected to carry out end-to-end workflows, producing complete artifacts from high-level user instructions. To meet enterprise needs, frontier AI labs have developed agents that can construct …

OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation

2026-04-13 · Xiaomeng Hu, Yinger Zhang, Fei Huang, Jianhong Tu 외 arxiv

AI agents are expected to perform professional work across hundreds of occupational domains (from emergency department triage to nuclear reactor safety monitoring to customs import processing), yet existing benchmarks ca…

Response Generation

SENTINEL: A Multi-Level Formal Framework for Safety Evaluation of Foundation Model-based Embodied Agents

2025-10-14 · Simon Sinong Zhan, Yao Liu, Philip Wang, Zinan Wang 외 arxiv

We present SENTINEL, a framework for formally evaluating the physical safety of foundation model (FM)-based embodied agents. SENTINEL is the first to provide multi-level safety evaluation across semantic interpretation, …

The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems

2026-02-19 · Leon Staufer, Kevin Feng, Kevin Wei, Luke Bailey 외 arxiv

Agentic AI systems are increasingly capable of performing professional and personal tasks with limited human involvement. However, tracking these developments is difficult because the AI agent ecosystem is complex, rapid…