paper-with-me

홈 › Papers

PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents

2026-08-04 · Shuhan Xue, Zixin Ding, Yichen Shen, Yinjie Wang, Zhenfei Yin, Yingcheng Wu, Yuxin Chen, Mengdi Wang, Ling Yang hf

Recursive self-improvement requires agents to turn accumulated experience into better future behavior. Personal AI agents offer a concrete setting for studying this capability because they retain preferences, task histories, tool routines, and learned skills across sessions. Yet whether retained experience actually improves them over time has not been systematically tested. We introduce PAST-Bench, a benchmark designed to isolate this question. Each agent runs through ordered sequences of fresh-session tasks under matched conditions that turn retained experience on and off. It spans 26 scenarios and 204 episodes across memory, procedural reuse, information gathering, and update. We report both later-task gains and whether those gains follow the intended save, retrieve, and update pathway. Across seven base models and four agent frameworks, improvement is real but uneven across capabilities. Agents with the same headline gain can differ markedly in whether that gain is supported by evidence of the intended pathway. Guided by these findings, we develop Hermes+, which extends Hermes with five targeted interventions across stages of the agent loop. Hermes+ raises the average gain from retained experience and provides clearer pathway evidence, with its strongest improvement on tasks requiring outdated state to be replaced, although the effect remains capability- and model-dependent. Together, PAST-Bench and Hermes+ provide an evaluation and diagnostic foundation for studying how persistent agents can progress from retaining experience to systematically improving through it. Code: https://github.com/Gen-Verse/PAST-Bench

📄 PDF Abstract BibTeX arXiv:2608.04003

Code (3)

Aaron617/agent-arXiv-daily ★ 10
Gen-Verse/PAST-Bench ★ 6
Tavish9/awesome-daily-AI-arxiv ★ 113

Similar Papers 제목 키워드 기반

Fibonacci-Driven Recursive Ensembles: Algorithms, Convergence, and Learning Dynamics

2026-01-03 · Ernest Fokoué arxiv

This paper develops the algorithmic and dynamical foundations of recursive ensemble learning driven by Fibonacci-type update flows. In contrast with classical boosting Freund and Schapire (1997); Friedman (2001), where t…

Ensemble Learning

Benchmarking Recursive-Collapse Warning Claims Under Matched False-Positive Control

2026-05-29 · David Mullett arxiv

Recursive systems can enter collapse-like regimes -- self-reinforcing amplification, persistent recursion, and narrowing diversity that mask accelerating internal degradation -- before overt failure becomes visible. We i…

Escher-Loop: Mutual Evolution by Closed-Loop Self-Referential Optimization

2026-04-25 · Ziyang Liu, Xinyan Guo, Xuchen Wei, Han Hao 외 arxiv

While recent autonomous agents demonstrate impressive capabilities, they predominantly rely on manually scripted workflows and handcrafted heuristics, inherently limiting their potential for open-ended improvement. To ad…

Emotion-Gradient Metacognitive RSI (Part I): Theoretical Foundations and Single-Agent Architecture

2025-05-12 · Rintaro Ando

We present the Emotion-Gradient Metacognitive Recursive Self-Improvement (EG-MRSI) framework, a novel architecture that integrates introspective metacognition, emotion-based intrinsic motivation, and recursive self-modif…

Informativeness

RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement

2026-07-28 · Fanqing Meng, Lingxiao Du, Qiguang Chen, Ziqi Zhao 외 arxiv

Recursive self-improvement requires turning evidence of model failures into better models. Data-centric post-training research entails diagnosing capability gaps, designing and validating training-data strategies, and le…

Question Answering