paper-with-me

홈 › Papers

The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators

2026-06-24 · Alex Iacob, Andrej Jovanović, William F. Shen, Daniel Burkhardt, Meghdad Kurmanji, Nurbek Tastan, Lorenzo Sani, Niccolò Alberto Elia Venanzi, Ambroise Odonnat, Zeyu Cao, Bill Marino, Xinchi Qiu, Nicholas D. Lane arxiv

Self-improving agents are state-of-the-art (SOTA) on agentic coding benchmarks and have recently been extended to general domains. However, their search methods generally assume a stationary evaluation criterion: a fixed verifier, benchmark, or labeled dataset that remains valid as the agent improves. This ignores a central feature of evolution: species adapt as their environments change with them. We aim to bring the same principle to recursive self-improvement, making evaluation part of the improvement loop and opening search to evolving evaluators, adversarial objectives, and dynamic utilities that may surpass static benchmarks. We introduce the Red Queen Godel Machine (RQGM), an evolutionary framework for recursive self-improvement under non-stationary utilities. The RQGM makes this possible through controlled utility evolution: search is organized into epochs with a fixed within-epoch evaluation criterion, while the utility can be updated at epoch boundaries, so self-improvement guarantees hold per epoch as the objective evolves across them. We begin by showing that even on verifiable coding tasks, the RQGM improves test pass rate over the prior SOTA by adding a complementary agent-as-a-judge code-review signal. This signal is cheaper and the RQGM uses 1.35x-1.72x fewer tokens. We then turn to scientific paper writing and reviewing, and Olympiad-level proof writing and grading, where the RQGM improves performance over prior self-improving agents: co-evolved writers reach 1.78x-1.86x higher acceptance rates under a diverse agent-as-a-judge panel, while co-evolved graders reach 9% higher ground-truth accuracy. In paper reviewing, the strongest baseline reviewer over-accepts AI-generated papers at up to 1.91x the human rate. The RQGM corrects this by introducing an adversarial objective that discovers reviewers equally stringent on AI and human work.

📄 PDF Abstract BibTeX arXiv:2606.26294

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers

2025-08-20 · Ziyang Luo, Zhiqi Shen, Wenzhuo Yang, Zirui Zhao 외 arxiv

The Model Context Protocol has emerged as a transformative standard for connecting large language models to external data sources and tools, rapidly gaining adoption across major AI providers and development platforms. H…

Evaluation-Driven Development of LLM Agents: A Process Model and Reference Architecture

2024-11-21 · Boming Xia, Qinghua Lu, Liming Zhu, Zhenchang Xing 외

Large Language Models (LLMs) have enabled the emergence of LLM agents: autonomous systems capable of achieving under-specified goals and adapting post-deployment, often without explicit code or model changes. Evaluating …

test driven development

The evolution of queen control over worker reproduction in the social Hymenoptera

2017-02-16

A trademark of eusocial insect species is reproductive division of labor, in which workers forego their own reproduction while the queen produces almost all offspring. The presence of the queen is key for maintaining soc…

GSAR: Goal-State-Anchor Rewards for Mobile GUI Agents with Self-Evolving Data Synthesis

2026-08-24 · Long Zhang, Yuhan Chen, Chaoran Zhang, Wanxia Cao 외 arxiv

Vision-Language Models (VLMs) based GUI agents stand to benefit significantly from online reinforcement learning (RL). However, their training is bottlenecked by two fundamental issues: current data synthesis methods for…

Reinforcement Learning

Queen-Bee Agents: A BeeSpec-Centered Architecture for Governed Enterprise MCP Orchestration

2026-06-04 · Dutao Zhang, Liaotian arxiv

Enterprise agent systems increasingly need to connect large language models to private tools, internal knowledge, and Model Context Protocol (MCP) interfaces. In this setting, raw task capability is insufficient: organiz…