paper-with-me

홈 › Papers

Beyond F1: Evaluating Coverage and Failure Recovery in AI Model Security Scanners

2026-08-27 · Qianlong Lan, Vinothini Pandurangan, Anuj Kaul, Indranil Sanyal arxiv

Static scanners are increasingly used to identify executable or otherwise unsafe content in machine- learning artifacts, yet conventional evaluation metrics characterize only cases where a scanner yields a usable security judgment. We evaluate ModelScan, ModelAudit, and Fickling using a controlled, artifact-backed benchmark on a synthetic corpus of 170 Pickle and PyTorch focused artifacts across 145 specimen families, 135 of which have binary security ground truth and 10 of which are intentionally malformed without labels. We explicitly distinguish non-N/A coverage, analysis completion, definitive security decisions, non-security findings, and unsupported outcomes. On labeled families, ModelAudit produced definitive security decisions for all 135 families (100%), Fickling for 110 (81.5%), and ModelScan for 67 (49.6%). Conditional on making a definitive judgment, ModelScan achieved 100% precision, recall, and F1. Fickling identified no unique true- positive families beyond those found by the combination of ModelAudit and ModelScan. Furthermore, for the 48 malicious families where ModelScan failed to complete its analysis, both ModelAudit and Fickling generated detections consistent with ground truth. These findings underscore the need to separate judgment accuracy from judgment availability, as well as incremental detection coverage from tool-level redundancy.

📄 PDF Abstract BibTeX arXiv:2608.27424

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SeClaw: Spec-Driven Security Task Synthesis for Evaluating Autonomous Agents

2026-06-01 · Hao Cheng, Changtao Miao, Tianle Song, Yin Wu 외 arxiv

Autonomous LLM agents increasingly operate in stateful environments where they access tools, files, memory, and external services. While such capabilities enable complex real-world workflows, they also introduce security…

Benchmarking Vision-Language-Action Models on SO-101: Failure and Recovery Analysis

2026-06-07 · Yi Yu, Xinchuan Qiu arxiv

Vision-Language-Action (VLA) models have demonstrated strong generalization in robotic manipulation, yet existing evaluations are primarily conducted in simulation or on expensive robotic platforms, leaving their robustn…

REPAIR-Bench: A Benchmark for Robot Error Perception And Interaction Recovery

2026-06-29 · Giuliano Pioldi, Yashika Batra, Arman Ibrayeva, Yuanchen Bai 외 arxiv

Understanding how users perceive and respond to robot failures is essential for building robust and trustworthy robot systems. Prior work, however, (i) often treats failures as independent events, (ii) emphasizes binary …

Beyond End-to-End Success: Diagnosing Failures in Long-Horizon Security LLM Agents

2026-08-20 · Wei Shao, Chongzhou Fang, Zuxiong Tan, Zequan Liang 외 arxiv

Long-horizon security LLM agents must carry information and decisions across many dependent interactions, where later actions often depend on services, state, or access discovered much earlier. This makes final task succ…

Decision Making

Artificial Intelligence Safety and Cybersecurity: a Timeline of AI Failures

2016-10-25 · Roman V. Yampolskiy, M. S. Spellchecker

In this work, we present and analyze reported failures of artificially intelligent systems and extrapolate our analysis to future AIs. We suggest that both the frequency and the seriousness of future AI failures will ste…