paper-with-me

홈 › Papers

When Skills Don't Help: A Negative Result on Procedural Knowledge for Tool-Grounded Agents in Offensive Cybersecurity

2026-05-19 · Samuel Jacob Chacko, James Hugglestone, Chashi Mahiul Islam, Xiuwen Liu arxiv

Agent Skills, structured packages of procedural knowledge loaded into an LLM agent at inference time, are widely reported to improve task pass rates by an average of 16.2~percentage points across diverse domains. Yet the same benchmarks show wide variance, with 16 of 84 tasks suffering negative deltas when Skills are introduced. The community has not yet articulated a clean mechanism for \emph{when} Skills help and when they are merely redundant overhead. We re-analyze a recently published 180-run controlled study of an MCP-grounded autonomous Capture-the-Flag (CTF) agent under four documentation conditions of increasing richness (591, 12865, 17253, and 36001 tokens) and show that these conditions correspond almost exactly to a No-Skills, Experiential-Skills, Curated-Skills, and Comprehensive-Skills ablation. In offensive cybersecurity, a domain not deeply covered by existing Skills benchmarks, the marginal benefit of Skills collapses. The spread between the no-Skills and full-Skills conditions is only 8.9~pp ($p = 0.71$, $χ^2$; $p = 0.25$, Cochran--Armitage trend test; five of six pairwise Cohen's $h$ values fall below the $0.2$ small-effect threshold). We argue that the missing variable is \emph{environment-feedback bandwidth}. When an agent's tool layer returns strict, schema-validated, low-latency observations, the environment itself supplies the procedural correction signal that Skills are normally needed to provide. As a result, the marginal benefit of curated Skills diminishes substantially, and, in some cases (e.g., our timing side-channel setting), actively degrades performance. We articulate a falsifiable hypothesis, sketch its design implications for compound AI systems, and will release the reanalysis pipeline to support replication.

📄 PDF Abstract BibTeX arXiv:2605.20023

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Demystifying Agent Skills: Why They Work-Until They Don't

2026-08-14 · Zhiyuan Jiang, Fangrui Huang, Hanwen Xing, Xander Wu 외 arxiv

Skills have emerged as a practical and effective approach for enhancing LLM agents at inference time through structured packages of knowledge. However, existing evaluations largely measure whether skills improve aggregat…

From Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skills

2026-05-22 · Zisu Huang, Jingwen Xu, Yifan Yang, Ziyang Gong 외 arxiv

Language agents increasingly improve by reusing \emph{skills} -- structured procedural artifacts distilled from past experience. In particular, \emph{domain-level} and \emph{model-generated} skills are especially promisi…

Managing Procedural Memory in LLM Agents: Control, Adaptation, and Evaluation

2026-06-22 · Julia Belikova, Rauf Parchiev, Evgeny Egorov, Grigorii Davydenko 외 hf

Procedural memory is increasingly used to improve LLM agents on recurring workplace tasks, yet its ability to produce reusable skills remains poorly understood. We introduce AFTER, a benchmark of 382 realistic enterprise…

SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills

2026-05-22 · Yingtie Lei, Zhongwei Wan, Jiankun Zhang, Samiul Alam 외 arxiv

Large language model (LLM) agents accumulate rich episodic trajectories while solving real-world tasks, but it remains unclear whether such experience can be distilled into reusable procedural skills. We introduce SkillE…

SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

2026-02-13 · Xiangyi Li, Yimin Liu, Wenbo Chen, Bingran You 외 arxiv

Agent Skills are structured packages of procedural knowledge that augment large language model (LLM) agents at inference time. Despite rapid adoption, there is no standard way to measure whether they actually help. We pr…