Probing In-Context Learning: Impact of Task Complexity and Model Architecture on Generalization and Efficiency
We investigate in-context learning (ICL) through a meticulous experimental framework that systematically varies task complexity and model architecture. Extending beyond the linear regression baseline, we introduce Gaussian kernel regression and nonlinear dynamical system tasks, which emphasize temporal and recursive reasoning. We evaluate four distinct models: a GPT2-style Transformer, a Transformer with FlashAttention mechanism, a convolutional Hyena-based model, and the Mamba state-space model. Each model is trained from scratch on synthetic datasets and assessed for generalization during testing. Our findings highlight that model architecture significantly shapes ICL performance. The standard Transformer demonstrates robust performance across diverse tasks, while Mamba excels in temporally structured dynamics. Hyena effectively captures long-range dependencies but shows higher variance early in training, and FlashAttention offers computational efficiency but is more sensitive in low-data regimes. Further analysis uncovers locality-induced shortcuts in Gaussian kernel tasks, enhanced nonlinear separability through input range scaling, and the critical role of curriculum learning in mastering high-dimensional tasks.
Code (1)
Tasks
Computational EfficiencyIn-Context LearningMambaregressionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Probing Linguistic Features of Sentence-Level Representations in Neural Relation Extraction
Despite the recent progress, little is known about the features captured by state-of-the-art neural relation extraction (RE) models. Common methods encode the source sentence, conditioned on the entity mentions, before c…
RelationRelation ExtractionSentencePareto Probing: Trading Off Accuracy for Complexity
The question of how to probe contextual word representations for linguistic structure in a way that is both principled and useful has seen significant attention recently in the NLP literature. In our contribution to this…
ARCDependency ParsingLearning Site-Specific Probing Beams for Fast mmWave Beam Alignment
Beam alignment - the process of finding an optimal directional beam pair - is a challenging procedure crucial to millimeter wave (mmWave) communication systems. We propose a novel beam alignment method that learns a site…
Domain-Informed Probing of wav2vec 2.0 Embeddings for Phonetic Features
In recent years large transformer model architectures have become available which provide a novel means of generating high-quality vector representations of speech audio. These transformers make use of an attention mecha…
Speaker Verificationspeech-recognitionSpeech RecognitionA Matter of Framing: The Impact of Linguistic Formalism on Probing Results
Deep pre-trained contextualized encoders like BERT (Delvin et al., 2019) demonstrate remarkable performance on a range of downstream tasks. A recent line of research in probing investigates the linguistic knowledge impli…