paper-with-me

Papers

Marti-5: A Mathematical Model of "Self in the World" as a First Step Toward Self-Awareness

2025-12-05 · Igor Pivovarov, Sergey Shumsky arxiv

The existence of 'what' and 'where' pathways of information processing in the brain was proposed almost 30 years ago, but there is still a lack of a clear mathematical model that could show how these pathways work together. We propose a biologically inspired mathematical model that uses this idea to identify and separate the self from the environment and then build and use a self-model for better predictions. This is a model of neocortical columns governed by the basal ganglia to make predictions and choose the next action, where some columns act as 'what' columns and others act as 'where' columns. Based on this model, we present a reinforcement learning agent that learns purposeful behavior in a virtual environment. We evaluate the agent on the Atari games Pong and Breakout, where it successfully learns to play. We conclude that the ability to separate the self from the environment gives advantages to the agent and therefore such a model could appear in living organisms during evolution. We propose Self-Awareness Principle 1: the ability to separate the self from the world is a necessary but insufficient condition for self-awareness.

📄 PDF Abstract BibTeX arXiv:2512.10985

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningAtari Games

Similar Papers 제목 키워드 기반

Semimartingale price systems in models with transaction costs beyond efficient friction

2020-01-09 · Christoph Kühn, Alexander Molitor

A standing assumption in the literature on proportional transaction costs is efficient friction. Together with robust no free lunch with vanishing risk, it rules out strategies of infinite variation, as they usually appe…

Friction

Information on trajectories: martingales and random times

2026-08-20 · Akshay Balsubramani arxiv

Accounting for information flow on the path space of trajectories of a nonnegative martingale yields exact variational identities for it, even at arbitrary random times. This recovers the widely used classical concentrat…

No Arbitrage in Continuous Financial Markets

2020-02-12

We derive integral tests for the existence and absence of arbitrage in a financial market with one risky asset which is either modeled as stochastic exponential of an Ito process or a positive diffusion with Markov switc…

Full-Step-DPO: Self-Supervised Preference Optimization with Step-wise Rewards for Mathematical Reasoning

2025-02-20 · Huimin Xu, Xin Mao, Feng-Lin Li, Xiaobao Wu 외

Direct Preference Optimization (DPO) often struggles with long-chain mathematical reasoning. Existing approaches, such as Step-DPO, typically improve this by focusing on the first erroneous step in the reasoning chain. H…

Mathematical Reasoning

S$^3$c-Math: Spontaneous Step-level Self-correction Makes Large Language Models Better Mathematical Reasoners

2024-09-03 · Yuchen Yan, Jin Jiang, Yang Liu, Yixin Cao 외

Self-correction is a novel method that can stimulate the potential reasoning abilities of large language models (LLMs). It involves detecting and correcting errors during the inference process when LLMs solve reasoning p…

GSM8KMathMathematical Reasoning