paper-with-me

Papers

Understanding In-Context Learning via Supportive Pretraining Data

2023-06-26 · Xiaochuang Han, Daniel Simig, Todor Mihaylov, Yulia Tsvetkov, Asli Celikyilmaz, Tianlu Wang

In-context learning (ICL) improves language models' performance on a variety of NLP tasks by simply demonstrating a handful of examples at inference time. It is not well understood why ICL ability emerges, as the model has never been specifically trained on such demonstrations. Unlike prior work that explores implicit mechanisms behind ICL, we study ICL via investigating the pretraining data. Specifically, we first adapt an iterative, gradient-based approach to find a small subset of pretraining data that supports ICL. We observe that a continued pretraining on this small subset significantly improves the model's ICL ability, by up to 18%. We then compare the supportive subset constrastively with random subsets of pretraining data and discover: (1) The supportive pretraining data to ICL do not have a higher domain relevance to downstream tasks. (2) The supportive pretraining data have a higher mass of rarely occurring, long-tail tokens. (3) The supportive pretraining data are challenging examples where the information gain from long-range context is below average, indicating learning to incorporate difficult long-range context encourages ICL. Our work takes a first step towards understanding ICL via analyzing instance-level pretraining data. Our insights have a potential to enhance the ICL ability of language models by actively guiding the construction of pretraining data in the future.

📄 PDF Abstract BibTeX arXiv:2306.15091

Code (0)

등록된 구현이 없습니다.

Tasks

In-Context Learning

Similar Papers 제목 키워드 기반

RedditESS: A Mental Health Social Support Interaction Dataset -- Understanding Effective Social Support to Refine AI-Driven Support Tools

2025-03-27 · Zeyad Alghamdi, Tharindu Kumarage, Garima Agrawal, Mansooreh Karami 외

Effective mental health support is crucial for alleviating psychological distress. While large language model (LLM)-based assistants have shown promise in mental health interventions, existing research often defines "eff…

Language ModelingLanguage ModellingLarge Language Model

Multi-Step Knowledge Interaction Analysis via Rank-2 Subspace Disentanglement

2025-11-03 · Sekh Mainul Islam, Pepa Atanasova, Isabelle Augenstein arxiv

Natural Language Explanations (NLEs) describe how Large Language Models (LLMs) make decisions by drawing on external Context Knowledge (CK) and Parametric Knowledge (PK). Understanding the interaction between these sourc…

Long-Term Feature Banks for Detailed Video Understanding

2018-12-12 · CVPR 2019 6 · Chao-yuan Wu, Christoph Feichtenhofer, Haoqi Fan, Kaiming He 외

To understand the world, we humans constantly need to relate the present to the past, and put events in context. In this paper, we enable existing video models to do the same. We propose a long-term feature bank---suppor…

Action ClassificationAction RecognitionEgocentric Activity RecognitionVideo Understanding

From Domain Understanding to Design Readiness: a playbook for GenAI-supported learning in Software Engineering

2026-03-31 · Rafal Wlodarski arxiv

Software engineering courses often require rapid upskilling in supporting knowledge areas such as domain understanding and modeling methods. We report an experience from a two-week milestone in a master's course where 29…

Incongruent Positivity: When Miscalibrated Positivity Undermines Online Supportive Conversations

2025-09-12 · Leen Almajed, Abeer ALdayel arxiv

In emotionally supportive conversations, well-intended positivity can sometimes misfire, leading to responses that feel dismissive, minimizing, or unrealistically optimistic. We examine this phenomenon of incongruent pos…