paper-with-me

홈 › Papers

Memory Efficient Neural Processes via Constant Memory Attention Block

2023-05-23 · Leo Feng, Frederick Tung, Hossein Hajimirsadeghi, Yoshua Bengio, Mohamed Osama Ahmed

Neural Processes (NPs) are popular meta-learning methods for efficiently modelling predictive uncertainty. Recent state-of-the-art methods, however, leverage expensive attention mechanisms, limiting their applications, particularly in low-resource settings. In this work, we propose Constant Memory Attentive Neural Processes (CMANPs), an NP variant that only requires constant memory. To do so, we first propose an efficient update operation for Cross Attention. Leveraging the update operation, we propose Constant Memory Attention Block (CMAB), a novel attention block that (i) is permutation invariant, (ii) computes its output in constant memory, and (iii) performs constant computation updates. Finally, building on CMAB, we detail Constant Memory Attentive Neural Processes. Empirically, we show CMANPs achieve state-of-the-art results on popular NP benchmarks while being significantly more memory efficient than prior methods.

📄 PDF Abstract BibTeX arXiv:2305.14567

Code (1)

borealisai/constant-memory-anp 공식 구현 pytorch

Tasks

Meta-Learning

Similar Papers 제목 키워드 기반

Constant Memory Attention Block

2023-06-21 · Leo Feng, Frederick Tung, Hossein Hajimirsadeghi, Yoshua Bengio 외

Modern foundation model architectures rely on attention mechanisms to effectively capture context. However, these methods require linear or quadratic memory in terms of the number of inputs/datapoints, limiting their app…

Point Processes

SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer

2025-09-29 · Junsong Chen, Yuyang Zhao, Jincheng Yu, Ruihang Chu 외 arxiv

We introduce SANA-Video, a small diffusion model that can efficiently generate videos up to 720x1280 resolution and minute-length duration. SANA-Video synthesizes high-resolution, high-quality and long videos with strong…

Video GenerationVideo Alignment

Echo: KV-Cache-Free Associative Recall with Spectral Koopman Operators

2026-05-07 · Anupama Sridhar, Alexander Johansen arxiv

Long chain-of-thought reasoning and agentic tool-calling produce traces spanning tens of thousands of tokens, yet Transformer KV caches grow linearly with sequence length, creating a memory bottleneck on commodity hardwa…

Tensor Memory: Fixed-Size Recurrent State for Long-Horizon Transformers

2026-05-26 · Kabir Swain, Sijie Han, Daniel Karl I. Weidele, Mauro Martino 외 arxiv

Transformers process images and videos by flattening space and time into long token sequences. While attention and KV caching preserve past features, their memory grows with sequence length and they lack an explicit, per…

xLSTM: Extended Long Short-Term Memory

2024-05-07 · Maximilian Beck, Korbinian Pöppel, Markus Spanring, Andreas Auer 외

In the 1990s, the constant error carousel and gating were introduced as the central ideas of the Long Short-Term Memory (LSTM). Since then, LSTMs have stood the test of time and contributed to numerous deep learning succ…

Language ModelingLanguage ModellingState Space Models