paper-with-me

홈 › Papers

The Dual-Route Model of Induction

2025-04-03 · Sheridan Feucht, Eric Todd, Byron Wallace, David Bau

Prior work on in-context copying has shown the existence of induction heads, which attend to and promote individual tokens during copying. In this work we introduce a new type of induction head: concept-level induction heads, which copy entire lexical units instead of individual tokens. Concept induction heads learn to attend to the ends of multi-token words throughout training, working in parallel with token-level induction heads to copy meaningful text. We show that these heads are responsible for semantic tasks like word-level translation, whereas token induction heads are vital for tasks that can only be done verbatim, like copying nonsense tokens. These two "routes" operate independently: in fact, we show that ablation of token induction heads causes models to paraphrase where they would otherwise copy verbatim. In light of these findings, we argue that although token induction heads are vital for specific tasks, concept induction heads may be more broadly relevant for in-context learning.

📄 PDF Abstract BibTeX arXiv:2504.03022

Code (0)

등록된 구현이 없습니다.

Tasks

In-Context Learningmodel

Similar Papers 제목 키워드 기반

Directional Routing in Transformers

2026-03-16 · Kevin Taylor arxiv

We introduce directional routing, a lightweight mechanism that gives each transformer attention head learned suppression directions controlled by a shared router, at 3.9% parameter cost. We train a 433M-parameter model a…

Duality Regularization for Unsupervised Bilingual Lexicon Induction

2019-09-03 · Xuefeng Bai, Yue Zhang, Hailong Cao, Tiejun Zhao

Unsupervised bilingual lexicon induction naturally exhibits duality, which results from symmetry in back-translation. For example, EN-IT and IT-EN induction can be mutually primal and dual problems. Current state-of-the-…

Bilingual Lexicon InductionTranslation

Universal Response and Emergence of Induction in LLMs

2024-11-11 · Niclas Luick

While induction is considered a key mechanism for in-context learning in LLMs, understanding its precise circuit decomposition beyond toy models remains elusive. Here, we study the emergence of induction behavior within …

In-Context Learning

Classification-Based Self-Learning for Weakly Supervised Bilingual Lexicon Induction

2020-07-01 · ACL 2020 6 · Mladen Karan, Ivan Vuli{\'c}, Anna Korhonen, Goran Glava{\v{s}}

Effective projection-based cross-lingual word embedding (CLWE) induction critically relies on the iterative self-learning procedure. It gradually expands the initial small seed dictionary to learn improved cross-lingual …

Bilingual Lexicon InductionClassificationGeneral ClassificationSelf-Learning+1

RouterDC: Query-Based Router by Dual Contrastive Learning for Assembling Large Language Models

2024-09-30 · Shuhao Chen, Weisen Jiang, Baijiong Lin, James T. Kwok 외

Recent works show that assembling multiple off-the-shelf large language models (LLMs) can harness their complementary abilities. To achieve this, routing is a promising method, which learns a router to select the most su…

Contrastive Learning