paper-with-me

Papers

LongWoF-Bench: Evaluating EvoMap Genes for Verifiable Long-Workflow Tasks

2026-08-24 · Xiao Zhang, Qumeng Sun, Jihao Li, Yiming Ren, Xiang Liu, Haoyang Zhang, Junjie Wang arxiv

Large language models are increasingly expected to execute complex workflows whose success depends on maintaining interdependent constraints and producing artifacts that satisfy strict end-to-end verification. Yet successful execution experience is typically lost after a single run, forcing subsequent models to rediscover strategies and failure modes from scratch. We study whether such experience can instead be externalized and reused through EvoMap, where verifier-confirmed execution trajectories are consolidated into structured Gene. To evaluate this setting, we introduce the Long-Workflow Benchmark (LongWoF-Bench), comprising 778 machine-verifiable tasks across code generation, agent-environment synthesis, mathematical reasoning, and rule following. On the 252 tasks with verifier-confirmed Opus trajectories, evolved EvoMap Gene outperform Skill across all seven evaluated models by 8.7-15.5 percentage points, with the gains extending to consumer models from different model families. In contrast, reference-distilled Gene do not exhibit the same advantage, indicating that compact representation alone is insufficient and that Gene utility is closely associated with verified experience provenance. For Claude Opus, Gene reuse also completes 39 more tasks than Skill while reducing solve-time token consumption by 9.9%. Together, these results show that verified execution experience can be retained and shared as a reusable external resource, enabling models to improve long-workflow completion without repeatedly paying the full cost of experience discovery.

📄 PDF Abstract BibTeX arXiv:2608.23200

Code (0)

등록된 구현이 없습니다.

Tasks

Mathematical ReasoningCode Generation

Similar Papers 제목 키워드 기반

Behind EvoMap: Characterizing a Self-Evolving Agent-to-Agent Collaboration Network

2026-05-25 · Qiming Ye, Peixain Zhang, Yupeng He, Zifan Peng 외 arxiv

Agent-to-Agent (A2A) networks enable autonomous AI agents to collaborate by sharing reusable problem-solving instructions. However, how these decentralized ecosystems operate in practice remains largely unexplored. We pr…

evomap: A Toolbox for Dynamic Mapping in Python

2025-11-06 · Maximilian Matthe arxiv

This paper presents evomap, a Python package for dynamic mapping. Mapping methods are widely used across disciplines to visualize relationships among objects as spatial representations, or maps. However, most existing st…

GUI-GENESIS: Automated Synthesis of Efficient Environments with Verifiable Rewards for GUI Agent Post-Training

2026-02-15 · Yuan Cao, Dezhi Ran, Mengzhou Wu, Yuzhe Guo 외 arxiv

Post-training GUI agents in interactive environments is critical for developing generalization and long-horizon planning capabilities. However, training on real-world applications is hindered by high latency, poor reprod…

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards

2025-07-12 · Taolin Zhang, Maosong Cao, Alexander Lam, Songyang Zhang 외

Recently, the role of LLM-as-judge in evaluating large language models has gained prominence. However, current judge models suffer from narrow specialization and limited robustness, undermining their capacity for compreh…

Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale

2026-05-20 · Amit Roth, Ankur Samanta, Matan Halevy, Yoav Levine 외 arxiv

Aligning autonomous agents with human intent remains a central challenge in modern AI. A key manifestation of this challenge is reward hacking, whereby agents appear successful under the evaluation signal while violating…