paper-with-me

홈 › Papers

Wanting to be Understood

2025-04-09 · Chrisantha Fernando, Dylan Banarse, Simon Osindero

This paper explores an intrinsic motivation for mutual awareness, hypothesizing that humans possess a fundamental drive to understand and to be understood even in the absence of extrinsic rewards. Through simulations of the perceptual crossing paradigm, we explore the effect of various internal reward functions in reinforcement learning agents. The drive to understand is implemented as an active inference type artificial curiosity reward, whereas the drive to be understood is implemented through intrinsic rewards for imitation, influence/impressionability, and sub-reaction time anticipation of the other. Results indicate that while artificial curiosity alone does not lead to a preference for social interaction, rewards emphasizing reciprocal understanding successfully drive agents to prioritize interaction. We demonstrate that this intrinsic motivation can facilitate cooperation in tasks where only one agent receives extrinsic reward for the behaviour of the other.

📄 PDF Abstract BibTeX arXiv:2504.06611

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Mutual Wanting in Human--AI Interaction: Empirical Evidence from Large-Scale Analysis of GPT Model Transitions

2025-10-27 · HaoYang Shang, Xuan Liu arxiv

The rapid evolution of large language models (LLMs) creates complex bidirectional expectations between users and AI systems that are poorly understood. We introduce the concept of "mutual wanting" to analyze these expect…

Wanting to Be Understood Explains the Meta-Problem of Consciousness

2025-06-10 · Chrisantha Fernando, Dylan Banarse, Simon Osindero

Because we are highly motivated to be understood, we created public external representations -- mime, language, art -- to externalise our inner states. We argue that such external representations are a pre-condition for …

Can Differentiable Decision Trees Enable Interpretable Reward Learning from Human Feedback?

2023-06-22 · Akansha Kalra, Daniel S. Brown

Reinforcement Learning from Human Feedback (RLHF) has emerged as a popular paradigm for capturing human intent to alleviate the challenges of hand-crafting the reward values. Despite the increasing interest in RLHF, most…

Atari GamesDiagnostic

Optimization of a SSP's Header Bidding Strategy using Thompson Sampling

2018-07-09 · Grégoire Jauvion, Nicolas Grislain, Pascal Sielenou Dkengne, Aurélien Garivier 외

Over the last decade, digital media (web or app publishers) generalized the use of real time ad auctions to sell their ad spaces. Multiple auction platforms, also called Supply-Side Platforms (SSP), were created. Because…

Thompson Sampling

Resolution Dependent GAN Interpolation for Controllable Image Synthesis Between Domains

2020-10-11 · Justin N. M. Pinkney, Doron Adler

GANs can generate photo-realistic images from the domain of their training data. However, those wanting to use them for creative purposes often want to generate imagery from a truly novel domain, a task which GANs are in…

Image Generation