paper-with-me

Papers

Blue Teaming Function-Calling Agents

2026-01-14 · Greta Dolcetti, Giulio Zizzo, Sergio Maffeis arxiv

We present an experimental evaluation that assesses the robustness of four open source LLMs claiming function-calling capabilities against three different attacks, and we measure the effectiveness of eight different defences. Our results show how these models are not safe by default, and how the defences are not yet employable in real-world scenarios.

📄 PDF Abstract BibTeX arXiv:2601.09292

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SIRAJ: Diverse and Efficient Red-Teaming for LLM Agents via Distilled Structured Reasoning

2025-10-30 · Kaiwen Zhou, Ahmed Elgohary, A S M Iftekhar, Amin Saied arxiv

The ability of LLM agents to plan and invoke tools exposes them to new safety risks, making a comprehensive red-teaming system crucial for discovering vulnerabilities and ensuring their safe deployment. We present SIRAJ:…

Benchmarking LLM-Assisted Blue Teaming via Standardized Threat Hunting

2025-09-28 · Yuqiao Meng, Luoxi Tang, Feiyang Yu, Xi Li 외 arxiv

As cyber threats continue to grow in scale and sophistication, blue team defenders increasingly require advanced tools to proactively detect and mitigate risks. Large Language Models (LLMs) offer promising capabilities f…

The Promise and Peril of Artificial Intelligence -- Violet Teaming Offers a Balanced Path Forward

2023-08-28 · Alexander J. Titus, Adam H. Russell

Artificial intelligence (AI) promises immense benefits across sectors, yet also poses risks from dual-use potentials, biases, and unintended behaviors. This paper reviews emerging issues with opaque and uncontrollable AI…

EthicsPhilosophyRed Teaming

Geometric Red-Teaming for Robotic Manipulation

2025-09-15 · Divyam Goel, Yufei Wang, Tiancheng Wu, Guixiu Qiao 외 arxiv

Standard evaluation protocols in robotic manipulation typically assess policy performance over curated, in-distribution test sets, offering limited insight into how systems fail under plausible variation. We introduce Ge…

Attack Atlas: A Practitioner's Perspective on Challenges and Pitfalls in Red Teaming GenAI

2024-09-23 · Ambrish Rawat, Stefan Schoepf, Giulio Zizzo, Giandomenico Cornacchia 외

As generative AI, particularly large language models (LLMs), become increasingly integrated into production applications, new attack surfaces and vulnerabilities emerge and put a focus on adversarial threats in natural l…

Red Teaming