paper-with-me

홈 › Papers

Survival is the Only Reward: Sustainable Self-Training Through Environment-Mediated Selection

2026-01-18 · Jennifer Dodgson, Alfath Daryl Alhajir, Michael Joedhitya, Akira Rafhael Janson Pattirane, Surender Suresh Kumar, Joseph Lim, C. H. Peh, Adith Ramdas, Steven Zhang Zhexu arxiv

Self-training systems often degenerate due to the lack of an external criterion for judging data quality, leading to reward hacking and semantic drift. This paper provides a proof-of-concept system architecture for stable self-training under sparse external feedback and bounded memory, and empirically characterises its learning dynamics and failure modes. We introduce a self-training architecture in which learning is mediated exclusively by environmental viability, rather than by reward, objective functions, or externally defined fitness criteria. Candidate behaviours are executed under real resource constraints, and only those whose environmental effects both persist and preserve the possibility of future interaction are propagated. The environment does not provide semantic feedback, dense rewards, or task-specific supervision; selection operates solely through differential survival of behaviours as world-altering events, making proxy optimisation impossible and rendering reward-hacking evolutionarily unstable. Analysis of semantic dynamics shows that improvement arises primarily through the persistence of effective and repeatable strategies under a regime of consolidation and pruning, a paradigm we refer to as negative-space learning (NSL), and that models develop meta-learning strategies (such as deliberate experimental failure in order to elicit informative error messages) without explicit instruction. This work establishes that environment-grounded selection enables sustainable open-ended self-improvement, offering a viable path toward more robust and generalisable autonomous systems without reliance on human-curated data or complex reward shaping.

📄 PDF Abstract BibTeX arXiv:2601.12310

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Evolutionary Self-Replication as a Mechanism for Producing Artificial Intelligence

2021-09-16 · Samuel Schmidgall, Joseph Hays

Can reproduction alone in the context of survival produce intelligence in our machines? In this work, self-replication is explored as a mechanism for the emergence of intelligent behavior in modern learning environments.…

Atari Games

On Reward Function for Survival

2016-06-18 · Naoto Yoshida

Obtaining a survival strategy (policy) is one of the fundamental problems of biological agents. In this paper, we generalize the formulation of previous research related to the survival of an agent and we formulate the s…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

ESG-coherent risk measures for sustainable investing

2023-09-11 · Gabriele Torri, Rosella Giacometti, Darinka Dentcheva, Svetlozar T. Rachev 외

The growing interest in sustainable investing calls for an axiomatic approach to measures of risk and reward that focus not only on financial returns, but also on measures of environmental and social sustainability, i.e.…

Survival Instinct in Offline Reinforcement Learning

2023-06-05 · NeurIPS 2023 11

We present a novel observation about the behavior of offline reinforcement learning (RL) algorithms: on many benchmark datasets, offline RL can produce well-performing and safe policies even when trained with "wrong" rew…

Offline RLreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Cooperate or Collapse: Emergence of Sustainable Cooperation in a Society of LLM Agents

2024-04-25 · Giorgio Piatti, Zhijing Jin, Max Kleiman-Weiner, Bernhard Schölkopf 외

As AI systems pervade human life, ensuring that large language models (LLMs) make safe decisions remains a significant challenge. We introduce the Governance of the Commons Simulation (GovSim), a generative simulation pl…

Decision MakingSpecificity