paper-with-me

홈 › Papers

MAS-PromptBench: When Does Prompt Optimization Improve Multi-Agent LLM Systems?

2026-06-22 · Juyang Bai, Laixi Shi arxiv

Multi-agent systems (MAS) offer a scalable path forward for agentic AI, comprising multiple LLM-based agents, each assigned a system prompt and a position within a workflow that governs inter-agent coordination and output aggregation. System prompts thus form a critical and accessible optimization surface: they specify agents' roles and behaviors, enabling system-level improvements without model finetuning. Although prompt optimization has shown substantial potential for single LLMs, extending it to MAS poses distinct challenges, notably an exponentially growing search space. It remains unclear whether, when, and by how much prompt optimization improves MAS performance, and how sensitive such gains are to system configuration. In this work, we systematically study system-prompt optimization across a broad range of MAS setups varying in task, workflow, communication protocol, and team size, benchmarking two prompt optimizers that naturally extend state-of-the-art single-agent methods. The results reveal its potential to unlock significant gains while exposing open challenges, characterizing when and how much prompt optimization helps across diverse MAS settings.

📄 PDF Abstract BibTeX arXiv:2606.23664

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards Self-Robust LLMs: Intrinsic Prompt Noise Resistance via CoIPO

2026-02-09 · Xin Yang, Letian Li, Abudukelimu Wuerkaixi, Xuxin Cheng 외 arxiv

Large language models (LLMs) have demonstrated remarkable and steadily improving performance across a wide range of tasks. However, LLM performance may be highly sensitive to prompt variations especially in scenarios wit…

Contrastive Learning

PromptBench: A Unified Library for Evaluation of Large Language Models

2023-12-13 · Kaijie Zhu, Qinlin Zhao, Hao Chen, Jindong Wang 외

The evaluation of large language models (LLMs) is crucial to assess their performance and mitigate potential security risks. In this paper, we introduce PromptBench, a unified library to evaluate LLMs. It consists of sev…

Prompt Engineering

Contrastive Instruction Tuning

2024-02-17 · Tianyi Lorena Yan, Fei Wang, James Y. Huang, Wenxuan Zhou 외

Instruction tuning has been used as a promising approach to improve the performance of large language models (LLMs) on unseen tasks. However, current LLMs exhibit limited robustness to unseen instructions, generating inc…

Sentence

DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

2023-09-29 · Kaijie Zhu, Jiaao Chen, Jindong Wang, Neil Zhenqiang Gong 외

Large language models (LLMs) have achieved remarkable performance in various evaluation benchmarks. However, concerns are raised about potential data contamination in their considerable volume of training corpus. Moreove…

Logical Reasoning

TRPrompt: Bootstrapping Query-Aware Prompt Optimization from Textual Rewards

2025-07-24 · Andreea Nica, Ivan Zakazov, Nicolas Mario Baldwin, Saibo Geng 외 arxiv

Prompt optimization improves the reasoning abilities of large language models (LLMs) without requiring parameter updates to the target model. Following heuristic-based "Think step by step" approaches, the field has evolv…