paper-with-me

홈 › Papers

On Reward Function for Survival

2016-06-18 · Naoto Yoshida

Obtaining a survival strategy (policy) is one of the fundamental problems of biological agents. In this paper, we generalize the formulation of previous research related to the survival of an agent and we formulate the survival problem as a maximization of the multi-step survival probability in future time steps. We introduce a method for converting the maximization of multi-step survival probability into a classical reinforcement learning problem. Using this conversion, the reward function (negative temporal cost function) is expressed as the log of the temporal survival probability. And we show that the objective function of the reinforcement learning in this sense is proportional to the variational lower bound of the original problem. Finally, We empirically demonstrate that the agent learns survival behavior by using the reward function introduced in this paper.

📄 PDF Abstract BibTeX arXiv:1606.05767

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Addressing reward bias in Adversarial Imitation Learning with neutral reward functions

2020-09-20 · Rohit Jena, Siddharth Agrawal, Katia Sycara

Generative Adversarial Imitation Learning suffers from the fundamental problem of reward bias stemming from the choice of reward functions used in the algorithm. Different types of biases also affect different types of e…

Imitation Learning

Reward Engineering for Spatial Epidemic Simulations: A Reinforcement Learning Platform for Individual Behavioral Learning

2025-11-22 · Radman Rakhshandehroo, Daniel Coombs arxiv

We present ContagionRL, a Gymnasium-compatible reinforcement learning platform specifically designed for systematic reward engineering in spatial epidemic simulations. Unlike traditional agent-based models that rely on f…

Reinforcement Learning

Survival Multiarmed Bandits with Bootstrapping Methods

2024-10-21 · Peter Veroutis, Frédéric Godin

The Multiarmed Bandits (MAB) problem has been extensively studied and has seen many practical applications in a variety of fields. The Survival Multiarmed Bandits (S-MAB) open problem is an extension which constrains an …

Predator-prey survival pressure is sufficient to evolve swarming behaviors

2023-08-24 · Jianan Li, Liang Li, Shiyu Zhao

The comprehension of how local interactions arise in global collective behavior is of utmost importance in both biological and physical research. Traditional agent-based models often rely on static rules that fail to cap…

Diversityreinforcement-learningReinforcement Learning

medR: Reward Engineering for Clinical Offline Reinforcement Learning via Tri-Drive Potential Functions

2026-02-03 · Qianyi Xu, Gousia Habib, Feng Wu, Yanrui Du 외 arxiv

Reinforcement Learning (RL) offers a powerful framework for optimizing dynamic treatment regimes (DTRs). However, clinical RL is fundamentally bottlenecked by reward engineering: the challenge of defining signals that sa…

Reinforcement Learning