paper-with-me

홈 › Papers

How Benchmarks and Evaluation Protocols Shape Conclusions in Provenance-Based Intrusion Detection

2026-08-02 · Lorenzo Guerra, Thomas Chapuis, Guillaume Duc, Pavlo Mozharovskyi, Van-Tam Nguyen arxiv

Provenance-based intrusion detection systems (PIDS) frequently report strong performance, but the conclusions drawn from these results can be highly sensitive to benchmarking choices and evaluation protocols. We investigate this dependency by re-evaluating representative PIDS on public datasets that meet our audit, labeling, and calibration requirements. Focusing primarily on the audited DARPA TC E3 datasets, we apply a unified protocol with temporally separated test periods and validation-only checkpoint selection and threshold calibration, and ask which architectural claims are empirically supported. We find that alerting success and investigation utility can diverge sharply, as several systems surface attacks without providing enough process-level context to support forensic investigation. Across the four primary datasets, a simple allowlist built from executable names and paths observed during training matches or exceeds the selected learned baselines on key operating-point metrics, showing that comparable performance on these metrics is achievable using lexical novelty alone. Quantifying semantic signal quality through feature completeness and field entropy helps explain why several audited E3 datasets support alerting performance without reliably separating model architectures. In contrast, Theia provides the richest semantic signal and shows the clearest improvements in ranking and node-level recovery for our reference model. Overall, these findings reinforce the importance of interpreting architectural claims in PIDS together with the benchmark properties and evaluation protocol that produced them.

📄 PDF Abstract BibTeX arXiv:2608.01454

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ManipArena: Comprehensive Real-world Evaluation of Reasoning-Oriented Generalist Robot Manipulation

2026-03-30 · Yu Sun, Meng Cao, Yang Ping, Kaidong Zhang 외 arxiv

Vision-Language-Action (VLA) models and world-action models have emerged as central paradigms for general-purpose robotic intelligence, yet their empirical progress remains constrained by the absence of evaluation protoc…

Robot Manipulation

Provenance-Based Interpretation of Multi-Agent Information Analysis

2020-11-08 · Scott Friedman, Jeff Rye, David LaVergne, Dan Thomsen 외

Analytic software tools and workflows are increasing in capability, complexity, number, and scale, and the integrity of our workflows is as important as ever. Specifically, we must be able to inspect the process of analy…

Diversity

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents

2026-04-11 · Miles Q. Li, Benjamin C. M. Fung, Boyang Li, Heba Ismail 외 arxiv

The rapid deployment of LLM-based autonomous agents has introduced safety risks that extend far beyond traditional LLM concerns, prompting a proliferation of safety benchmarks since late 2023. However, these benchmarks h…

Data Provenance via Differential Auditing

2022-09-04 · Xin Mu, Ming Pang, Feida Zhu

Auditing Data Provenance (ADP), i.e., auditing if a certain piece of data has been used to train a machine learning model, is an important problem in data provenance. The feasibility of the task has been demonstrated by …

DeconDTN-Toolkit: A Library for Evaluation and Enhancement of Robustness to Provenance Shift

2026-05-11 · Yongsen Tan, Zhecheng Sheng, Xiruo Ding, Serguei V. S. Pakhomov 외 arxiv

Despite the burgeoning body of work on distribution shifts, provenance shift-where the relationship between data source and label changes at deployment-remains poorly understood and under-addressed. In this paper, we est…