paper-with-me

홈 › Papers

Co-LMLM: Continuous-Query Limited Memory Language Models

2026-07-08 · Yair Feldman, Linxi Zhao, Nathan Godey, Dongyoung Go, Yilun Hua, Kilian Q. Weinberger, Jennifer J. Sun, Yoav Artzi arxiv

Limited memory language models (LMLMs) externalize factual knowledge during pretraining to a knowledge base (KB), rather than memorizing it in their weights. During generation, the model then fetches knowledge from the KB as needed. This recently introduced paradigm provides multiple advantages, including knowledge control capabilities that remain beyond conventional LLMs. We propose continuous-query LMLM (CO-LMLM), where the KB pairs continuous keys with textual knowledge values, a significant departure from prior reliance on relational KB and queries. CO-LMLM generates flexible vector queries at minimal cost, while still integrating human-readable and attributable retrieved knowledge into its generation. We pair this design with an annotation pipeline that tags free-form factual spans in arbitrary text, removing prior work's restriction to Wikipedia. Across pretraining on Wikipedia and FineWeb-Edu and at multiple model scales, CO-LMLM outperforms prior LMLMs and vanilla LLMs in both perplexity and factual precision. At 360M scale, this includes lower perplexity than models pretrained on 40x more data, and SimpleQA-verified performance that is in line with gpt-4o-mini and higher than Claude Sonnet 4.5.

📄 PDF Abstract BibTeX arXiv:2607.07707

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Auditing Forgetting in Limited Memory Language Models

2026-07-01 · Arya Raeesi, Hanna Roed arxiv

Limited Memory Language Models (LMLMs) externalize factual knowledge to a database to enable deletion-based unlearning without retraining. Existing evaluations measure post-deletion correctness in aggregate and cannot te…

MLMLM: Link Prediction with Mean Likelihood Masked Language Model

2020-09-15 · Findings (ACL) 2021 8 · Louis Clouatre, Philippe Trempe, Amal Zouaq, Sarath Chandar

Knowledge Bases (KBs) are easy to query, verifiable, and interpretable. They however scale with man-hours and high-quality data. Masked Language Models (MLMs), such as BERT, scale with computing power as well as unstruct…

Language ModelingLanguage ModellingLink PredictionPrediction

Pre-training Large Memory Language Models with Internal and External Knowledge

2025-05-21 · Linxi Zhao, Sofian Zalouk, Christian K. Belardi, Justin Lovelace 외

Neural language models are black-boxes -- both linguistic patterns and factual knowledge are distributed across billions of opaque parameters. This entangled encoding makes it difficult to reliably inspect, verify, or up…

Memorization

Improving Large Molecular Language Model via Relation-aware Multimodal Collaboration

2026-01-18 · Jinyoung Park, Minseong Bae, Jeehye Na, Hyunwoo J. Kim arxiv

Large language models (LLMs) have demonstrated their instruction-following capabilities and achieved powerful performance on various tasks. Inspired by their success, recent works in the molecular domain have led to the …

Molecule Captioning

HingeMem: Boundary Guided Long-Term Memory with Query Adaptive Retrieval for Scalable Dialogues

2026-04-08 · Yijie Zhong, Yunfan Gao, Haofen Wang arxiv

Long-term memory is critical for dialogue systems that support continuous, sustainable, and personalized interactions. However, existing methods rely on continuous summarization or OpenIE-based graph construction paired …

Event SegmentationQuestion Answering