paper-with-me

홈 › Papers

General-Purpose In-Context Learning by Meta-Learning Transformers

2022-12-08 · Louis Kirsch, James Harrison, Jascha Sohl-Dickstein, Luke Metz

Modern machine learning requires system designers to specify aspects of the learning pipeline, such as losses, architectures, and optimizers. Meta-learning, or learning-to-learn, instead aims to learn those aspects, and promises to unlock greater capabilities with less manual effort. One particularly ambitious goal of meta-learning is to train general-purpose in-context learning algorithms from scratch, using only black-box models with minimal inductive bias. Such a model takes in training data, and produces test-set predictions across a wide range of problems, without any explicit definition of an inference model, training loss, or optimization algorithm. In this paper we show that Transformers and other black-box models can be meta-trained to act as general-purpose in-context learners. We characterize transitions between algorithms that generalize, algorithms that memorize, and algorithms that fail to meta-train at all, induced by changes in model size, number of tasks, and meta-optimization. We further show that the capabilities of meta-trained algorithms are bottlenecked by the accessible state size (memory) determining the next prediction, unlike standard models which are thought to be bottlenecked by parameter count. Finally, we propose practical interventions such as biasing the training distribution that improve the meta-training and meta-generalization of general-purpose in-context learning algorithms.

📄 PDF Abstract BibTeX arXiv:2212.04458

Code (1)

RobvanGastel/meta-in-context-learning jax

Tasks

In-Context LearningInductive BiasMeta-Learning

Methods 이 논문이 사용한 방법론

fail 설명 없음

Similar Papers 제목 키워드 기반

MetaMorph: Learning Universal Controllers with Transformers

2022-03-22 · ICLR 2022 4 · Agrim Gupta, Linxi Fan, Surya Ganguli, Li Fei-Fei

Multiple domains like vision, natural language, and audio are witnessing tremendous progress by leveraging Transformers for large scale pre-training followed by task specific fine tuning. In contrast, in robotics we prim…

Zero-shot Generalization

In-Context Learning through the Bayesian Prism

2023-06-08 · Madhur Panwar, Kabir Ahuja, Navin Goyal

In-context learning (ICL) is one of the surprising and useful features of large language models and subject of intense research. Recently, stylized meta-learning-like ICL setups have been devised that train transformers …

Bayesian InferenceIn-Context LearningInductive BiasLanguage Modelling+2

Transformers are almost optimal metalearners for linear classification

2025-10-22 · Roey Magen, Gal Vardi arxiv

Transformers have demonstrated impressive in-context learning (ICL) capabilities, raising the question of whether they can serve as metalearners that adapt to new tasks using only a small number of in-context examples, w…

Hierarchical Transformers are Efficient Meta-Reinforcement Learners

2024-02-09 · Gresa Shala, André Biedenkapp, Josif Grabocka

We introduce Hierarchical Transformers for Meta-Reinforcement Learning (HTrMRL), a powerful online meta-reinforcement learning approach. HTrMRL aims to address the challenge of enabling reinforcement learning agents to p…

Meta Reinforcement Learningreinforcement-learningReinforcement Learning

Meta-Learning Transformers to Improve In-Context Generalization

2025-07-07 · Lorenzo Braccaioli, Anna Vettoruzzo, Prabhant Singh, Joaquin Vanschoren 외

In-context learning enables transformer models to generalize to new tasks based solely on input prompts, without any need for weight updates. However, existing training paradigms typically rely on large, unstructured dat…

In-Context LearningMeta-Learning