paper-with-me

홈 › Papers

Dual Process Learning: Controlling Use of In-Context vs. In-Weights Strategies with Weight Forgetting

2024-05-28 · Suraj Anand, Michael A. Lepori, Jack Merullo, Ellie Pavlick

Language models have the ability to perform in-context learning (ICL), allowing them to flexibly adapt their behavior based on context. This contrasts with in-weights learning, where information is statically encoded in model parameters from iterated observations of the data. Despite this apparent ability to learn in-context, language models are known to struggle when faced with unseen or rarely seen tokens. Hence, we study $\textbf{structural in-context learning}$, which we define as the ability of a model to execute in-context learning on arbitrary tokens -- so called because the model must generalize on the basis of e.g. sentence structure or task structure, rather than semantic content encoded in token embeddings. An ideal model would be able to do both: flexibly deploy in-weights operations (in order to robustly accommodate ambiguous or unknown contexts using encoded semantic information) and structural in-context operations (in order to accommodate novel tokens). We study structural in-context algorithms in a simple part-of-speech setting using both practical and toy models. We find that active forgetting, a technique that was recently introduced to help models generalize to new languages, forces models to adopt structural in-context learning solutions. Finally, we introduce $\textbf{temporary forgetting}$, a straightforward extension of active forgetting that enables one to control how much a model relies on in-weights vs. in-context solutions. Importantly, temporary forgetting allows us to induce a $\textit{dual process strategy}$ where in-context and in-weights solutions coexist within a single model.

📄 PDF Abstract BibTeX arXiv:2406.00053

Code (0)

등록된 구현이 없습니다.

Tasks

In-Context Learning

Similar Papers 제목 키워드 기반

Breaking through the learning plateaus of in-context learning in Transformer

2023-09-12 · Jingwen Fu, Tao Yang, Yuwang Wang, Yan Lu 외

In-context learning, i.e., learning from context examples, is an impressive ability of Transformer. Training Transformers to possess this in-context learning skill is computationally intensive due to the occurrence of le…

In-Context LearningRepresentation Learning

Controlling Explanatory Heatmap Resolution and Semantics via Decomposition Depth

2016-03-21 · Sebastian Bach, Alexander Binder, Klaus-Robert Müller, Wojciech Samek

We present an application of the Layer-wise Relevance Propagation (LRP) algorithm to state of the art deep convolutional neural networks and Fisher Vector classifiers to compare the image perception and prediction strate…

Prediction

FADO: Feedback-Aware Double COntrolling Network for Emotional Support Conversation

2022-11-01 · Wei Peng, Ziyuan Qin, Yue Hu, Yuqiang Xie 외

Emotional Support Conversation (ESConv) aims to reduce help-seekers'emotional distress with the supportive strategy and response. It is essential for the supporter to select an appropriate strategy with the feedback of t…

Response Generation

Local Contrastive Editing of Gender Stereotypes

2024-10-23 · Marlene Lutz, Rochelle Choenni, Markus Strohmaier, Anne Lauscher

Stereotypical bias encoded in language models (LMs) poses a threat to safe language technology, yet our understanding of how bias manifests in the parameters of LMs remains incomplete. We introduce local contrastive edit…

Fine-Grained Controllable Text Generation Using Non-Residual Prompting

2021-11-16 · ACL ARR November 2021 11 · Anonymous

The introduction of immensely large Causal Language Models (CLMs) has rejuvenated the interest in open-ended text generation. However, controlling the generative process for these Transformer-based models is at large an …

DecoderText Generation