paper-with-me

Papers

SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior

2026-06-10 · Zhiyu Chen, Zihan Guo, Bo Huang, Bingwei Lu, Jianghao Lin, Yuanjian Zhou, Weinan Zhang arxiv

Agent Skills augment large language model (LLM) agents with procedural knowledge at inference time, but current benchmarks rarely distinguish what a Skill says from how it is organized. We study this distinction through Progressive Disclosure, where a concise root file points agents to supporting resources on demand, and compare it with a normalized flat baseline. We present SkillJuror, a framework for evaluating Skill writing paradigms through semantically controlled variants, matched multi-trial evaluations, and trajectory evidence while holding task knowledge fixed. In an 82-task SkillsBench study, Progressive Disclosure changes runtime behavior before aggregate outcomes: distinct Skill resources touched per trajectory rise from 1.18 to 3.85, and effective uptake events rise from 1.33 to 3.92. It also yields 17 additional verifier-passing trials out of 410 matched trials (+4.1%) over the normalized flat baseline. The benefit is task-dependent. Progressive Disclosure helps when supporting resources guide implementation, checking, or repair, but is weaker when success hinges on exact output conventions, numerical thresholds, or long artifact-generation pipelines. These results show that Skill organization is not mere presentation: it can change how agents search and apply procedural knowledge, while outcome gains depend on whether the exposed resources are actionable for the task. Code is available at https://github.com/zhiyuchen-ai/skill-juror.

📄 PDF Abstract BibTeX arXiv:2606.11543

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A simple method for joint evaluation of skill in directional forecasts of multiple variables

2024-02-02 · Thitithep Sitthiyot, Kanyarat Holasut

Forecasts for key macroeconomic variables are almost always made simultaneously by the same organizations, presented together, and used together in policy analyses and decision-makings. It is therefore important to know …

Behavioral Analysis of Vision-and-Language Navigation Agents

2023-07-20 · CVPR 2023 1 · Zijiao Yang, Arjun Majumdar, Stefan Lee

To be successful, Vision-and-Language Navigation (VLN) agents must be able to ground instructions to actions based on their surroundings. In this work, we develop a methodology to study agent behavior on a skill-specific…

Vision and Language Navigation

Counterfactual Trace Auditing of LLM Agent Skills

2026-05-12 · Xiaolin Zhou, Jinbo Liu, Li Li, Ryan A. Rossi 외 arxiv

Large Language Model agents are increasingly augmented with agent skills. Current evaluation methods for skills remain limited. Most deployed benchmarks report only pass rate before and after a skill is attached, treatin…

Why Solve It Twice? Hierarchical Accumulation of Skills for Transfer-Efficient ML Engineering

2026-06-29 · Yongbin Kim, Yashar Talebirad, Osmar R. Zaiane arxiv

ML engineering agents waste compute rediscovering known techniques because every competition is a cold start. We present HASTE, a hierarchical multi-agent system that organizes cross-competition knowledge into three scop…

Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks

2026-02-23 · David Schmotz, Luca Beurer-Kellner, Sahar Abdelnabi, Maksym Andriushchenko arxiv

LLM agents are evolving rapidly, powered by code execution, tools, and the recently introduced agent skills feature. Skills allow users to extend LLM applications with specialized third-party code, knowledge, and instruc…