paper-with-me

Papers

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops

2026-07-08 · Mingguang Chen, Licheng Wang, Bo Qu arxiv

AI systems increasingly participate in their own improvement: revising their outputs, adapting their own harnesses during deployment, training on data they generate, and, increasingly, conducting AI research itself. This literature is described under a vocabulary ("self-refine," "self-reward," "self-play," "self-evolve") that conflates fundamentally different ambitions. We survey 1,250 arXiv papers (2024-2026) along two axes: what the system improves -- its behavior in deployment, its policy through training, its evaluator, or the research process itself -- and the degree of loop closure (human-in-the-loop to fully closed). The taxonomy separates bounded self-refinement -- convergent, evaluable, and already industrial practice -- from open-ended recursive self-improvement (RSI), which remains bounded by grounding requirements, collapse dynamics, and compute constraints on every measured axis. Its distinctive feature is a dedicated category for self-evaluation: every improvement loop is a claim that some signal can substitute for human judgment. We survey the evaluator design space -- judges, process reward models, verifiers, rubrics, meta-evaluation -- order the signals into a verification hierarchy from formal verifiers (strongest) to intrinsic self-assessment (weakest), and observe that demonstrated self-improvement strength tracks this hierarchy, that its failure modes (self-confirming loops, model collapse, diversity collapse) follow from its violations, and that the "research direction-setting" bottleneck keeping humans in the loop sits at the top of that hierarchy. We connect the technical literature to the theory of RSI limits and to the safety and governance questions raised by frontier-lab accounts of closing the loop, and identify governance-grade measurement of self-improvement as the field's most underpopulated niche.

📄 PDF Abstract BibTeX arXiv:2607.07663

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Self-Aware Recursively Self-Improving Agents for Personal Singularity: A Goal-, Scope-, Tool-, and Benchmark-Driven Multi-Agent Architecture

2026-07-14 · Chengshuai Yang arxiv

Large language model (LLM) agents can plan, use tools, maintain memory, and execute long-horizon tasks. This paper proposes Self-Aware Recursively Self-Improving (SARSI) agents: governed agents that maintain a persistent…

Bounded Recursive Self-Improvement

2013-12-24 · E. Nivel, K. R. Thórisson, B. R. Steunebrink, H. Dindo 외

We have designed a machine that becomes increasingly better at behaving in underspecified circumstances, in a goal-directed way, on the job, by modeling itself and its environment as experience accumulates. Based on prin…

Scheduling

ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents

2026-09-15 · Shuhan Xue, Jianyuan Zhong, Ziyuan Nan, Wenbin Li 외 arxiv

We introduce and release ScienceBuddy, an interactive scientific research workspace that brings continually improving scientific agents into researchers' everyday workflows. ScienceBuddy supports researchers in carrying …

Reinforcement LearningContinual Learning

MetaRSI / RSI2: A Meta-Recursive Self-Improving System for Recursive Self-Improving Systems Themselves

2026-09-06 · Zihan Tan, Leixin Sun, Zitong Shi, Yitao Liu 외 arxiv

Recursive self-improvement (RSI) lets a system improve the model-building machinery from its own failures, so every later model inherits the gain. Yet RSI has been validated almost exclusively on coding and formal benchm…

AQuA: Recursively Self-Improving Quantitative Trading Research Agents

2026-08-13 · Jiacheng Guo, Suozhi Huang, Yunlong Gao, Zihao Li 외 arxiv

We study recursive self-improvement at the level of quantitative-investment research: whether an autonomous system can use evidence from earlier experiments to improve the hypotheses and candidates proposed in later iter…