paper-with-me

홈 › Papers

AgentCo-op: Retrieval-Based Synthesis of Interoperable Multi-Agent Workflows

2026-05-19 · Shuaike Shen, Wenduo Cheng, Shike Wang, Mingqian Ma, Jian Ma arxiv

Designing multi-agent workflows is especially difficult in open-ended scientific settings where tasks lack curated training sets, reliable scalar evaluation metrics, and standardized interfaces between existing tools and agents. We propose AgentCo-op, a retrieval-based synthesis framework that composes reusable skills, tools, and external agents into executable workflows through typed artifact handoffs, then applies bounded self-guided local repair to implicated components when execution evidence indicates failure. In two open-world genomics case studies, AgentCo-op composes independently developed scientific agents and external tool repositories into auditable workflows without redesigning them or running global topology search. It coordinates specialized agents for spatial transcriptomics and gene-set interpretation to enable collaborative discovery from spatial transcriptomics data, and builds a parallel workflow for cross-modality marker analysis on single-cell multiome data. AgentCo-op can also import a searched workflow as a structural prior and improve it by grounding nodes with retrieved components and applying local repair, showing that synthesis and search are complementary. On six coding, math, and question-answering benchmarks, AgentCo-op achieves the best result on four benchmarks and the best average score under a unified backbone setting, while consistently reducing per-task cost relative to multi-agent baselines. Together, these results suggest that retrieval-based synthesis can extend automated agentic workflow design beyond benchmark-optimized agent graphs to open-world workflows built from existing agents, tools, and typed artifacts.

📄 PDF Abstract BibTeX arXiv:2605.20425

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators

2026-05-09 · Aritra Mazumder, Shubhashis Roy Dipta, Nusrat Jahan Lia, Tanzila Khan 외 arxiv

Multi-agent systems achieve state-of-the-art outcomes through peer collaboration. However, when an agent in the pipeline silently drops a constraint, the system's final output may look correct even though the reasoning c…

AgentCompass: Towards Reliable Evaluation of Agentic Workflows in Production

2025-09-18 · NVJK Kartik, Garvit Sapra, Rishav Hada, Nikhil Pareek arxiv

With the growing adoption of Large Language Models (LLMs) in automating complex, multi-agent workflows, organizations face mounting risks from errors, emergent behaviors, and systemic failures that current evaluation met…

Continual Learning

AgentCourt: Simulating Court with Adversarial Evolvable Lawyer Agents

2024-08-15 · Guhong Chen, Liyang Fan, Zihan Gong, Nan Xie 외

In this paper, we present a simulation system called AgentCourt that simulates the entire courtroom process. The judge, plaintiff's lawyer, defense lawyer, and other participants are autonomous agents driven by large lan…

Textualized Agent-Style Reasoning for Complex Tasks by Multiple Round LLM Generation

2024-09-19 · Chen Liang, Zhifan Feng, Zihe Liu, Wenbin Jiang 외

Chain-of-thought prompting significantly boosts the reasoning ability of large language models but still faces three issues: hallucination problem, restricted interpretability, and uncontrollable generation. To address t…

Hallucination

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities

2026-07-15 · Kai Chen, Zichen Ding, Jiaye Ge, Shufan Jiang 외 arxiv

As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical. However, current evaluation pipelines remain highly fragmented and tightly coupled, hinderin…