paper-with-me

Papers

MIME: Mutual Information Minimisation Exploration

2020-01-16 · Haitao Xu, Brendan McCane, Lech Szymanski, Craig Atkinson

We show that reinforcement learning agents that learn by surprise (surprisal) get stuck at abrupt environmental transition boundaries because these transitions are difficult to learn. We propose a counter-intuitive solution that we call Mutual Information Minimising Exploration (MIME) where an agent learns a latent representation of the environment without trying to predict the future states. We show that our agent performs significantly better over sharp transition boundaries while matching the performance of surprisal driven agents elsewhere. In particular, we show state-of-the-art performance on difficult learning games such as Gravitar, Montezuma's Revenge and Doom.

📄 PDF Abstract BibTeX arXiv:2001.05636

Code (0)

등록된 구현이 없습니다.

Tasks

Montezuma's Revengereinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

DRMIME: Differentiable Mutual Information and Matrix Exponential for Multi-Resolution Image Registration

2020-01-27 · MIDL 2019 7 · Abhishek Nan, Matthew Tennant, Uriel Rubin, Nilanjan Ray

In this work, we present a novel unsupervised image registration algorithm. It is differentiable end-to-end and can be used for both multi-modal and mono-modal registration. This is done using mutual information (MI) as …

Image RegistrationUnsupervised Image Registration

Disambiguating Anthropomorphism and Anthropomimesis in Human-Robot Interaction

2026-02-10 · Minja Axelsson, Henry Shevlin arxiv

In this preliminary work, we offer an initial disambiguation of the theoretical concepts anthropomorphism and anthropomimesis in Human-Robot Interaction (HRI) and social robotics. We define anthropomorphism as users perc…

Mime: Mimicking Centralized Stochastic Algorithms in Federated Learning

2020-08-08 · Sai Praneeth Karimireddy, Martin Jaggi, Satyen Kale, Mehryar Mohri 외

Federated learning (FL) is a challenging setting for optimization due to the heterogeneity of the data across different clients which gives rise to the client drift phenomenon. In fact, obtaining an algorithm for FL whic…

Federated Learning

MIMEx: Intrinsic Rewards from Masked Input Modeling

2023-05-15 · NeurIPS 2023 11 · Toru Lin, Allan Jabri

Exploring in environments with high-dimensional observations is hard. One promising approach for exploration is to use intrinsic rewards, which often boils down to estimating "novelty" of states, transitions, or trajecto…

Prediction

Brain-Inspired Capture: Evidence-Driven Neuromimetic Perceptual Simulation for Visual Decoding

2026-04-20 · Feixue Shao, Guangze Shi, Xueyu Liu, Yongfei Wu 외 arxiv

Visual decoding of neurophysiological signals is a critical challenge for brain-computer interfaces (BCIs) and computational neuroscience. However, current approaches are often constrained by the systematic and stochasti…

Image Retrieval