paper-with-me

홈 › Papers

Learning by Distilling Context

2022-09-30 · Charlie Snell, Dan Klein, Ruiqi Zhong

Language models significantly benefit from context tokens, such as prompts or scratchpads. They perform better when prompted with informative instructions, and they acquire new reasoning capabilities by generating a scratch-pad before predicting the final answers. However, they do not \textit{internalize} these performance gains, which disappear when the context tokens are gone. Our work proposes to apply context distillation so that a language model can improve itself by internalizing these gains. Concretely, given a synthetic unlabeled input for the target task, we condition the model on `[instructions] + [task-input]'' to predict [scratch-pad] + [final answer]''; then we fine-tune the same model to predict its own [final answer]'' conditioned on the [task-input]'', without seeing the [instructions]'' or using the `[scratch-pad]''. We show that context distillation is a general method to train language models, and it can effectively internalize 3 types of training signals. First, it can internalize abstract task instructions and explanations, so we can iteratively update the model parameters with new instructions and overwrite old ones. Second, it can internalize step-by-step reasoning for complex tasks (e.g., 8-digit addition), and such a newly acquired capability proves to be useful for other downstream tasks. Finally, it can internalize concrete training examples, and it outperforms directly learning with gradient descent by 9\% on the SPIDER Text-to-SQL dataset; furthermore, combining context distillation operations can internalize more training examples than the context window size allows.

📄 PDF Abstract BibTeX arXiv:2209.15189

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModellingText to SQLText-To-SQL

Similar Papers 제목 키워드 기반

Whispering Context: Distilling Syntax and Semantics for Long Speech Transcripts

2025-08-18 · Duygu Altinok arxiv

ASR systems often struggle with maintaining syntactic and semantic accuracy in long audio transcripts, impacting tasks like Named Entity Recognition (NER), capitalization, and punctuation. We propose a novel approach tha…

Context-aware Difference Distilling for Multi-change Captioning

2024-05-31 · Yunbin Tu, Liang Li, Li Su, Zheng-Jun Zha 외

Multi-change captioning aims to describe complex and coupled changes within an image pair in natural language. Compared with single-change captioning, this task requires the model to have higher-level cognition ability t…

AllDecoder

Distilling Large Language Models for Network Active Queue Management

2025-01-28 · Deol Satish, Shiva Raj Pokhrel, Jonathan Kua, Anwar Walid

The growing complexity of network traffic and demand for ultra-low latency communication require smarter packet traffic management. Existing Deep Learning-based queuing approaches struggle with dynamic network scenarios …

Few-Shot LearningManagement

Improving Topic Modeling by Distilling Soft Labels from Language Models

2026-02-20 · Raymond Li, Amirhossein Abaskohi, Chuyuan Li, Gabriel Murray 외 arxiv

Traditional neural topic models are typically optimized by reconstructing the document's Bag-of-Words (BoW) representations, overlooking contextual information and struggling with data sparsity. In this work, we introduc…

Topic Models

Distilling Spectral Graph for Object-Context Aware Open-Vocabulary Semantic Segmentation

2024-11-26 · CVPR 2025 1 · Chanyoung Kim, Dayun Ju, Woojung Han, Ming-Hsuan Yang 외

Open-Vocabulary Semantic Segmentation (OVSS) has advanced with recent vision-language models (VLMs), enabling segmentation beyond predefined categories through various learning schemes. Notably, training-free methods off…

ObjectOpen Vocabulary Semantic SegmentationOpen-Vocabulary Semantic SegmentationSemantic Segmentation