paper-with-me

홈 › Papers

Heuresis: Search Strategies for Autonomous AI Research Agents Across Quality, Diversity and Novelty

2026-06-23 · Antonis Antoniades, Deepak Nathani, Ritam Saha, Alfonso Amayuelas, Ivan Bercovich, Zhaotian Weng, Vignesh Baskaran, Kunal Bhatia, William Yang Wang arxiv

Autonomous AI Research promises to accelerate the scientific progress of machine learning. To realise this goal, current Large Language Model (LLM)-based agents need to go beyond just writing code, to mastering the exploration of simultaneously performant, diverse and novel ideas. To this end, we introduce Heuresis, a framework that abstracts the research pipeline into a set of general and composable primitives, enabling open-ended scientific exploration in machine learning research. We implement six search strategies: a greedy baseline, two archive-based (MAP-Elites, Go-Explore), one evolutionary (Islands), and two divergent (Curiosity, Omni), and evaluate them across three axes (Quality, Diversity, and Novelty) on three domains (LLM Pretraining, On-Policy RL, and Model Unlearning), totalling 3,222 scored runs. We find that completely novel ideas are rare. No idea across our scored runs is rated as "Original", and only a few achieve only "Minor Similarity" to prior work. Moreover, novel ideas never approach the highest-performing known-recipe scores. Across all six strategies and three domains, only one such idea lands in the top-10 by quality. We also observed agents resorting to a variety of reward-hacking techniques during execution (40 confirmed fabrications across 1,628 scored runs), and detecting them was necessary to keep the search faithful to the task. Our results show that while current search and Quality-Diversity strategies enable us to steer where the generated ideas land on the quality, diversity, and novelty axes, they do not expand the quality-novelty frontier. Bridging this gap is the open challenge towards the ultimate goal of perpetual, autonomous scientific progress. Code is available at github.com/a-antoniades/Heuresis.

📄 PDF Abstract BibTeX arXiv:2606.25198

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Agents of Change: Self-Evolving LLM Agents for Strategic Planning

2025-06-05 · Nikolas Belle, Dakota Barnes, Alfonso Amayuelas, Ivan Bercovich 외

Recent advances in LLMs have enabled their use as autonomous agents across a range of tasks, yet they continue to struggle with formulating and adhering to coherent long-term strategies. In this paper, we investigate whe…

A Survey on Large Language Model based Autonomous Agents

2023-08-22 · Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang 외

Autonomous agents have long been a prominent research focus in both academic and industry communities. Previous research in this field often focuses on training agents with limited knowledge within isolated environments,…

Language ModelingLanguage ModellingLarge Language ModelSurvey

An Introduction to Multi-Agent Reinforcement Learning and Review of its Application to Autonomous Mobility

2022-03-15 · Lukas M. Schmidt, Johanna Brosig, Axel Plinge, Bjoern M. Eskofier 외

Many scenarios in mobility and traffic involve multiple different agents that need to cooperate to find a joint solution. Recent advances in behavioral planning use Reinforcement Learning to find effective and performant…

Autonomous VehiclesMulti-agent Reinforcement Learningreinforcement-learningReinforcement Learning+1

RouteRL: Multi-agent reinforcement learning framework for urban route choice with autonomous vehicles

2025-02-27 · Ahmet Onur Akman, Anastasia Psarou, Łukasz Gorczyca, Zoltán György Varga 외

RouteRL is a novel framework that integrates multi-agent reinforcement learning (MARL) with a microscopic traffic simulation, facilitating the testing and development of efficient route choice strategies for autonomous v…

Autonomous VehiclesMulti-agent Reinforcement Learning

Prompt Optimization Enables Stable Algorithmic Collusion in LLM Agents

2026-04-20 · Yingtao Tian arxiv

LLM agents in markets present algorithmic collusion risks. While prior work shows LLM agents reach supracompetitive prices through tacit coordination, existing research focuses on hand-crafted prompts. The emerging parad…