paper-with-me

홈 › Papers

MoME: Mixture-of-Memory Embeddings for Context-Aware Sparse Lookup

2026-09-14 · Muchen Li, Leonid Sigal, Renjie Liao hf

Scaling large language models efficiently has motivated sparse capacity mechanisms such as Mixture-of-Experts and, more recently, conditional memory: token-indexed embedding tables that augment the backbone with cheap parametric lookups. Existing memory-embedding methods retrieve via a deterministic function of the surface form, which collapses different contextual senses of the same token (e.g., python the language vs. the animal) into a single fixed entry. We introduce Mixture of Memory Embeddings (MoME), a context-aware memory mechanism that replaces each token's single memory row with a mixture of M slots and uses a learned gate over the hidden state to choose which slots to read at each position. In controlled pretraining experiments across nanochat, Llama-3/MobileLLM, and Qwen3 backbones, MoME improves over Value Embedding, Bigram, and STEM baselines in iso-parameter and iso-training-FLOP settings, shows a more promising memory-size scaling trend at sub-billion scale, and remains efficient in training and inference. Qualitative routing analyses on polysemous tokens further suggest that the learned mixture exhibits a degree of semantic interpretability, dispatching the same surface token to distinct memory slots under different senses.

📄 PDF Abstract BibTeX arXiv:2609.15126

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Amory: Building Coherent Narrative-Driven Agent Memory through Agentic Reasoning

2026-01-09 · Yue Zhou, Xiaobo Guo, Belhassen Bayar, Srinivasan H. Sengamedu arxiv

Long-term conversational agents face a fundamental scalability challenge as interactions extend over time: repeatedly processing entire conversation histories becomes computationally prohibitive. Current approaches attem…

Exploring Speaker Diarization with Mixture of Experts

2025-06-17 · Gaobin Yang, Maokui He, Shutong Niu, Ruoyu Wang 외

In this paper, we propose a novel neural speaker diarization system using memory-aware multi-speaker embedding with sequence-to-sequence architecture (NSD-MS2S), which integrates a memory-aware multi-speaker embedding mo…

Mixture-of-Expertsspeaker-diarizationSpeaker Diarization

MoME: Mixture of Visual Language Medical Experts for Medical Imaging Segmentation

2025-10-30 · Arghavan Rezvani, Xiangyi Yan, Anthony T. Wu, Kun Han 외 arxiv

In this study, we propose MoME, a Mixture of Visual Language Medical Experts, for Medical Image Segmentation. MoME adapts the successful Mixture of Experts (MoE) paradigm, widely used in Large Language Models (LLMs), for…

Medical Image Segmentation

Where Should Optimizer State Live? Tiered State Allocation for Memory-Efficient Mixture-of-Experts Training

2026-07-21 · Nuemaan Malik hf

Optimizer state is the largest single line item in the memory budget of mixture-of-experts (MoE) training. On a 6.78B-parameter MoE language model AdamW keeps 50.6 GB of first and second moments to update 12.6 GB of bflo…

Context-Aware Mixture-of-Experts Inference on CXL-Enabled GPU-NDP Systems

2025-12-04 · Zehao Fan, Zhenyu Liu, Yunzhen Liu, Yayue Hou 외 arxiv

Mixture-of-Experts (MoE) models scale large language models through conditional computation, but inference becomes memory-bound once expert weights exceed the capacity of GPU memory. In this case, weights must be offload…