paper-with-me

홈 › Papers

Understanding In-Context Learning Beyond Transformers: An Investigation of State Space and Hybrid Architectures

2025-10-27 · Shenran Wang, Timothy Tin-Long Tse, Jian Zhu arxiv

We perform in-depth evaluations of in-context learning (ICL) on state-of-the-art transformer, state-space, and hybrid large language models over two categories of knowledge-based ICL tasks. Using a combination of behavioral probing and intervention-based methods, we have discovered that, while LLMs of different architectures can behave similarly in task performance, their internals could remain different. We discover that function vectors (FVs) responsible for ICL are primarily located in the self-attention and Mamba layers, and speculate that Mamba2 uses a different mechanism from FVs to perform ICL. FVs are more important for ICL involving parametric knowledge retrieval, but not for contextual knowledge understanding. Our work contributes to a more nuanced understanding across architectures and task types. Methodologically, our approach also highlights the importance of combining both behavioural and mechanistic analyses to investigate LLM capabilities.

📄 PDF Abstract BibTeX arXiv:2510.23006

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CodeSSM: Towards State Space Models for Code Understanding

2025-05-02 · Shweta Verma, Abhinav Anand, Mira Mezini

Although transformers are widely used for various code-specific tasks, they have some significant limitations. In this paper, we investigate State Space Models (SSMs) as a potential alternative to transformers for code u…

Clone DetectionLanguage ModelingLanguage ModellingMasked Language Modeling+3

Vision Transformers for Computer Go

2023-09-22 · Amani Sagri, Tristan Cazenave, Jérôme Arjonilla, Abdallah Saffidine

Motivated by the success of transformers in various fields, such as language understanding and image analysis, this investigation explores their application in the context of the game of Go. In particular, our study focu…

Game of Go

ViT-5: Vision Transformers for The Mid-2020s

2026-02-08 · Feng Wang, Sucheng Ren, Tiezheng Zhang, Predrag Neskovic 외 arxiv

This work presents a systematic investigation into modernizing Vision Transformer backbones by leveraging architectural advancements from the past five years. While preserving the canonical Attention-FFN structure, we co…

Representation LearningSpatial Reasoning

Transformers Can Learn Posterior Predictive Distributions In-Context

2026-05-26 · Gyeonghun Kang, Changwoo J. Lee, Xiang Cheng arxiv

Prior-data fitted networks (PFNs) have recently emerged as a powerful approach for Bayesian prediction tasks, approximating the posterior predictive distribution (PPD) through in-context learning. Despite their strong em…

Interpreting Affine Recurrence Learning in GPT-style Transformers

2024-10-22 · Samarth Bhargav, Alexander Gu

Understanding the internal mechanisms of GPT-style transformers, particularly their capacity to perform in-context learning (ICL), is critical for advancing AI alignment and interpretability. In-context learning allows t…

In-Context Learning