paper-with-me

Papers

Grounded Skill Synthesis from Code at Scale for Agentic Intelligence

2026-09-04 · Yongqi Tong, Pan Wang, Hang Wang, Jianshe Li, Xin Zhang, Jiang-Ming Yang, Wei Wu hf

Reusable skills give agents transferable procedural knowledge, making scalable acquisition essential for extending agents beyond prior experience. Existing methods face two limitations: trajectory-based synthesis requires interactions with specific environments, while document-derived skills may lack executable evidence and verification. Source code offers a complementary path: it requires no prior agent experience yet provides executable evidence for grounding abstractions. We present Code2Skill, a fully automated pipeline that transforms selected code units into implementation-anchored records of atomic operations, composite workflows, and recurring patterns, then verifies each record through source-body-blind reconstruction and source-aware comparison. Applied to 19,769 popular, actively maintained GitHub repositories, Code2Skill produces CodeSkillBank, a grounded bank of 1,006,822 accepted records with workflow, boundary, provenance, and source-evidence metadata. Across 72 protocol-matched evaluations covering nine model settings and eight benchmarks, models augmented with retrieved CodeSkillBank skills improve by 11.7% on average over matched baselines and outperform them in 57 cases. Under a unified downstream interface, Code2Skill also outperforms trajectory-derived skill banks on all seven shared benchmarks, showing that repository-derived skills can provide useful procedural knowledge before agents accumulate sufficient interaction experience. Skills synthesized from tested AI-generated code achieve a 93.50% pass rate, compared with 93.00% for human-written code, providing initial evidence that the pipeline can expand with the growing volume of AI-generated software. Overall, Code2Skill transforms procedural knowledge embedded in repositories into grounded, verifiable, and transferable agent skills.

📄 PDF Abstract BibTeX arXiv:2609.05571

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MidTool: Mid-training Data Synthesis for Agentic Tool Use

2026-08-20 · Fengqing Jiang, Yite Wang, Boyi Liu, Zhaoyang Wang 외 arxiv

Mid-training is increasingly recognized as a critical stage for shaping the capabilities of large language models. Recent work has shown that targeted mid-training can strengthen reasoning-intensive abilities such as mat…

Reinforcement Learning

SoK: Agentic Skills -- Beyond Tool Use in LLM Agents

2026-02-24 · Yanna Jiang, Delong Li, Haiyu Deng, Baihe Ma 외 arxiv

Agentic systems increasingly rely on reusable procedural capabilities, \textit{a.k.a., agentic skills}, to execute long-horizon workflows reliably. These capabilities are callable modules that package procedural knowledg…

SkillCoach: Self-Evolving Rubrics for Evaluating and Enhancing Agentic Skill-Use

2026-07-02 · Jiayin Zhu, Kelong Mao, Yudong Guo, Dengbo He 외 arxiv

Skills are becoming a reusable operational layer for LLM agents, encoding SOPs, domain rules, tool workflows, scripts, and validation routines. In realistic skill repositories, overlapping skills make reliable skill-use …

Dr. RTL: Autonomous Agentic RTL Optimization through Tool-Grounded Self-Improvement

2026-04-16 · Wenji Fang, Yao Lu, Shang Liu, Jing Wang 외 arxiv

Recent advances in large language models (LLMs) have sparked growing interest in automatic RTL optimization for better performance, power, and area (PPA). However, existing methods are still far from realistic RTL optimi…

S1-DeepResearch: Beyond Search, Toward Real-World Long-Horizon Research Agents

2026-06-13 · Yao Dong, Xinglin Xiao, Liwei Dong, Xinlong Jin 외 arxiv

Deep research agents aim to solve complex knowledge-intensive tasks through long-horizon planning, evidence gathering, reasoning, and report generation. While recent progress in search agents has demonstrated strong capa…

Instruction FollowingInformation RetrievalQuestion Answering