paper-with-me

홈 › Papers

Transformer verbatim in-context retrieval across time and scale

2024-11-11 · Kristijan Armeni, Marko Pranjić, Senja Pollak

To predict upcoming text, language models must in some cases retrieve in-context information verbatim. In this report, we investigated how the ability of language models to retrieve arbitrary in-context nouns developed during training (across time) and as language models trained on the same dataset increase in size (across scale). We then asked whether learning of in-context retrieval correlates with learning of more challenging zero-shot benchmarks. Furthermore, inspired by semantic effects in human short-term memory, we evaluated the retrieval with respect to a major semantic component of target nouns, namely whether they denote a concrete or abstract entity, as rated by humans. We show that verbatim in-context retrieval developed in a sudden transition early in the training process, after about 1% of the training tokens. This was observed across model sizes (from 14M and up to 12B parameters), and the transition occurred slightly later for the two smallest models. We further found that the development of verbatim in-context retrieval is positively correlated with the learning of zero-shot benchmarks. Around the transition point, all models showed the advantage of retrieving concrete nouns as opposed to abstract nouns. In all but two smallest models, the advantage dissipated away toward the end of training.

📄 PDF Abstract BibTeX arXiv:2411.07075

Code (1)

kristijanarmeni/verbatim-memory-in-nlms 공식 구현 pytorch

Tasks

Retrieval

Similar Papers 제목 키워드 기반

Characterizing Verbatim Short-Term Memory in Neural Language Models

2022-10-24 · Kristijan Armeni, Christopher Honey, Tal Linzen

When a language model is trained to predict natural language sequences, its prediction at each moment depends on a representation of prior context. What kind of information about the prior context can language models ret…

Language ModellingRetrieval

Structured Distillation for Personalized Agent Memory: 11x Token Reduction with Retrieval Preservation

2026-03-13 · Sydney Lewis arxiv

Long conversations with an AI agent create a simple problem for one user: the history is useful, but carrying it verbatim is expensive. We study personalized agent memory: one user's conversation history with an agent, d…

Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation Systems

2024-02-27 · Zhenting Qi, HANLIN ZHANG, Eric Xing, Sham Kakade 외

Retrieval-Augmented Generation (RAG) improves pre-trained models by incorporating external knowledge at test time to enable customized adaptation. We study the risk of datastore leakage in Retrieval-In-Context RAG Langua…

Instruction FollowingRAGRetrievalRetrieval-augmented Generation

In-Context Molecular Property Prediction with LLMs: A Blinding Study on Memorization and Knowledge Conflicts

2026-03-26 · Matthias Busch, Marius Tacke, Sviatlana V. Lamaka, Mikhail L. Zheludkevich 외 arxiv

The capabilities of large language models (LLMs) have expanded beyond natural language processing to scientific prediction tasks, including molecular property prediction. However, their effectiveness in in-context learni…

Molecular Property Prediction

Panini: Continual Learning in Token Space via Structured Memory

2026-02-16 · Shreyas Rajesh, Pavan Holur, Mehmet Yigit Turali, Chenda Duan 외 arxiv

Language models are increasingly used to reason over content they were not trained on, such as new documents, evolving knowledge, and user-specific data. A common approach is retrieval-augmented generation (RAG), which s…

Continual Learning