paper-with-me

홈 › Papers

Re-examining learning linear functions in context

2024-11-18 · Omar Naim, Guilhem Fouilhé, Nicholas Asher

In-context learning (ICL) has emerged as a powerful paradigm for easily adapting Large Language Models (LLMs) to various tasks. However, our understanding of how ICL works remains limited. We explore a simple model of ICL in a controlled setup with synthetic training data to investigate ICL of univariate linear functions. We experiment with a range of GPT-2-like transformer models trained from scratch. Our findings challenge the prevailing narrative that transformers adopt algorithmic approaches like linear regression to learn a linear function in-context. These models fail to generalize beyond their training distribution, highlighting fundamental limitations in their capacity to infer abstract task structures. Our experiments lead us to propose a mathematically precise hypothesis of what the model might be learning.

📄 PDF Abstract BibTeX arXiv:2411.11465

Code (0)

등록된 구현이 없습니다.

Tasks

In-Context Learning

Methods 이 논문이 사용한 방법론

Linear Regression Linear Regression is a method for modelling a relationship between a dependent variable and independent variables. These models can be fit with numerous approaches. The most…
ADOPT Please enter a description about the method here

Similar Papers 제목 키워드 기반

Axioms for AI Alignment from Human Feedback

2024-05-23 · Luise Ge, Daniel Halpern, Evi Micha, Ariel D. Procaccia 외

In the context of reinforcement learning from human feedback (RLHF), the reward function is generally derived from maximum likelihood estimation of a random utility model based on pairwise comparisons made by humans. The…

How does Lipschitz Regularization Influence GAN Training?

2018-11-23 · ECCV 2020 8 · Yipeng Qin, Niloy Mitra, Peter Wonka

Despite the success of Lipschitz regularization in stabilizing GAN training, the exact reason of its effectiveness remains poorly understood. The direct effect of $K$-Lipschitz regularization is to restrict the $L2$-norm…

Polyhedral Complex Derivation from Piecewise Trilinear Networks

2024-02-16 · Jin-Hwa Kim

Recent advancements in visualizing deep neural networks provide insights into their structures and mesh extraction from Continuous Piecewise Affine (CPWA) functions. Meanwhile, developments in neural surface representati…

Representation Learning

Protecting Privacy in Classifiers by Token Manipulation

2024-07-01 · Re'em Harel, Yair Elboher, Yuval Pinter

Using language models as a remote service entails sending private information to an untrusted provider. In addition, potential eavesdroppers can intercept the messages, thereby exposing the information. In this work, we …

text-classificationText Classification

Re-Examining Linear Embeddings for High-Dimensional Bayesian Optimization

2020-01-31 · NeurIPS 2020 12 · Benjamin Letham, Roberto Calandra, Akshara Rai, Eytan Bakshy

Bayesian optimization (BO) is a popular approach to optimize expensive-to-evaluate black-box functions. A significant challenge in BO is to scale to high-dimensional parameter spaces while retaining sample efficiency. A …

Bayesian OptimizationMisconceptionsVocal Bursts Intensity Prediction