paper-with-me

Papers

Environment-Grounded Multi-Agent Workflow for Autonomous Penetration Testing

2026-03-25 · Michael Somma, Markus Großpointner, Paul Zabalegui, Eppu Heilimo, Branka Stojanović arxiv

The increasing complexity and interconnectivity of digital infrastructures make scalable and reliable security assessment methods essential. Robotic systems represent a particularly important class of operational technology, as modern robots are highly networked cyber-physical systems deployed in domains such as industrial automation, logistics, and autonomous services. This paper explores the use of large language models for automated penetration testing in robotic environments. We propose an environment-grounded multi-agent architecture tailored to Robotics-based systems. The approach dynamically constructs a shared graph-based memory during execution that captures the observable system state, including network topology, communication channels, vulnerabilities, and attempted exploits. This enables structured automation while maintaining traceability and effective context management throughout the testing process. Evaluated across multiple iterations within a specialized robotics Capture-the-Flag scenario (ROS/ROS2), the system demonstrated high reliability, successfully completing the challenge in 100\% of test runs (n=5). This performance significantly exceeds literature benchmarks while maintaining the traceability and human oversight required by frameworks like the EU AI Act.

📄 PDF Abstract BibTeX arXiv:2603.24221

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?

2026-08-11 · Mizanur Rahman, Mohammed Saidul Islam, Ridwan Mahbub, Md Tahmid Rahman Laskar 외 hf

Real-world data science involves long-horizon workflows that span data wrangling, exploration, modeling, visualization, and validation, and require coordinated use of tools such as notebooks, IDEs, terminals, browsers, a…

DeepAnalyze: Agentic Large Language Models for Autonomous Data Science

2025-10-19 · Shaolei Zhang, Ju Fan, Meihao Fan, Guoliang Li 외 arxiv

Autonomous data science, from raw data sources to analyst-grade deep research reports, has been a long-standing challenge, and is now becoming feasible with the emergence of powerful large language models (LLMs). Recent …

Question Answering

PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments

2026-05-04 · Ruoqi Liu, Imran Q. Mohiuddin, Austin J. Schoeffler, Kavita Renduchintala 외 arxiv

We introduce PhysicianBench, a benchmark for evaluating LLM agents on physician tasks grounded in real clinical setting within electronic health record (EHR) environments. Existing medical agent benchmarks primarily focu…

Socratic agents for autonomous scientific discovery in high-dimensional physical systems

2026-06-25 · Xianrui Zeng, Pengfei Liu, Yirui Zang, Yang Shen 외 arxiv

The automation of scientific discovery has reached an inflection point. While AI systems now operate instruments, optimize parameters and generate hypotheses, most remain procedural: they execute workflows fixed by human…

CutVerse: A Compositional GUI Agents Benchmark for Media Post-Production Editing

2026-05-19 · Haobo Hu, Xiangwu Guo, Zhiheng Chen, Difei Gao 외 arxiv

While GUI agents have made significant progress in web navigation and basic operating system tasks, their capabilities in professional creative workflows remain largely underexplored. To bridge this gap, we introduce Cut…