paper-with-me

홈 › Papers

Kinaema: a recurrent sequence model for memory and pose in motion

2025-10-23 · Mert Bulent Sariyildiz, Philippe Weinzaepfel, Guillaume Bono, Gianluca Monaci, Christian Wolf arxiv

One key aspect of spatially aware robots is the ability to "find their bearings", ie. to correctly situate themselves in previously seen spaces. In this work, we focus on this particular scenario of continuous robotics operations, where information observed before an actual episode start is exploited to optimize efficiency. We introduce a new model, Kinaema, and agent, capable of integrating a stream of visual observations while moving in a potentially large scene, and upon request, processing a query image and predicting the relative position of the shown space with respect to its current position. Our model does not explicitly store an observation history, therefore does not have hard constraints on context length. It maintains an implicit latent memory, which is updated by a transformer in a recurrent way, compressing the history of sensor readings into a compact representation. We evaluate the impact of this model in a new downstream task we call "Mem-Nav". We show that our large-capacity recurrent model maintains a useful representation of the scene, navigates to goals observed before the actual episode start, and is computationally efficient, in particular compared to classical transformers with attention over an observation history.

📄 PDF Abstract BibTeX arXiv:2510.20261

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RPM-Net: Recurrent Prediction of Motion and Parts from Point Cloud

2020-06-26 · Zihao Yan, Ruizhen Hu, Xingguang Yan, Luanmin Chen 외

We introduce RPM-Net, a deep learning-based approach which simultaneously infers movable parts and hallucinates their motions from a single, un-segmented, and possibly partial, 3D point cloud shape. RPM-Net is a novel Re…

DecoderSemantic Segmentation

VideoSSM: Autoregressive Long Video Generation with Hybrid State-Space Memory

2025-12-04 · Yifei Yu, Xiaoshan Wu, Xinting Hu, Tao Hu 외 arxiv

Autoregressive (AR) diffusion enables streaming, interactive long-video generation by producing frames causally, yet maintaining coherence over minute-scale horizons remains challenging due to accumulated errors, motion …

Video Generation

Recurrent Attention Unit

2018-10-30 · Guoqiang Zhong, Guohua Yue, Xiao Ling

Recurrent Neural Network (RNN) has been successfully applied in many sequence learning problems. Such as handwriting recognition, image description, natural language processing and video motion analysis. After years of d…

General ClassificationHandwriting Recognitionimage-classificationImage Classification+5

Learning Video Object Segmentation with Visual Memory

2017-04-19 · ICCV 2017 10 · Pavel Tokmakov, Karteek Alahari, Cordelia Schmid

This paper addresses the task of segmenting moving objects in unconstrained videos. We introduce a novel two-stream neural network with an explicit memory module to achieve this. The two streams of the network encode spa…

Motion SegmentationObjectSemantic SegmentationUnsupervised Video Object Segmentation+2

Memory-Augmented Temporal Dynamic Learning for Action Recognition

2019-04-30 · Yuan Yuan, Dong Wang, Qi. Wang

Human actions captured in video sequences contain two crucial factors for action recognition, i.e., visual appearance and motion dynamics. To model these two aspects, Convolutional and Recurrent Neural Networks (CNNs and…

Action RecognitionTemporal Action Localization