paper-with-me

Papers

When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models

2026-03-22 · Abhinaba Basu arxiv

We identify a routing paradox in hybrid sequence models: content-based routing - deciding which tokens deserve expensive attention - requires pairwise computation, and this requirement is inescapable. Through 20+ controlled experiments across three tasks, multiple scales (200K to 1.4B parameters), and 15+ routing mechanisms, we map the routing landscape exhaustively. Every system that achieves high routing precision does so through pairwise token comparison. Every mechanism that avoids pairwise computation fails: recurrent models (Mamba-1.4B: 29%), memory banks (12%), bandits (0.7-3.6%), contrastive pretraining (1.6%), and 12 other approaches all cluster at 1-29%. Routing needs two ingredients: (1) per-token representations with bidirectional context and (2) pairwise token comparison. Bidirectional Mamba (O(n)) + pairwise comparison achieves 99.5%; replacing the full pairwise router with rank-1 projection improves this to 99.7%. Adding one bidirectional layer to frozen Pythia-1B recovers 99.4% routing. Six different O(n) preprocessing mechanisms (bidirectional Mamba, Perceiver inducing points, causal attention with E2E training, sparse attention, bidirectional attention, rank-1 projection) all succeed; global mean pooling (1.9%) and Fourier mixing (0.9%) fail. The routing signal occupies a ~34-dimensional latent subspace, invisible to cosine similarity. Non-learned indices (Bloom filter: 90.9%; BM25: 82.7%) bypass the bottleneck for exact/keyword matching. Combining O(n) bidirectional Mamba with rank-1 pairwise projection yields 99.7% routing at linear inference cost.

📄 PDF Abstract BibTeX arXiv:2603.20997

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

When Does Routing Become Interpretable? Causal Probes on Block Attention Residuals

2026-06-11 · Aydin Javadov arxiv

Block Attention Residuals (Block AttnRes) by replace fixed additive residuals with a learned softmax over earlier depth-source representations, surfacing cross-layer routing as an inspectable tensor in the forward pass. …

All Routes Lead to Collapse

2026-06-21 · K. R. Balasubramanian arxiv

Attention sinks, representation collapse, and norm stratification are treated as transformer-specific pathologies. We show they are not specific to attention: they are what content-based routing does under a fixed simila…

When Does Explicit View Routing Work? A Controlled Study of Multi-View Graph-Text Alignment

2026-07-29 · Xiao Yue, Guangzhi Qu arxiv

Graph-text retrieval typically maps a graph and its description to a single embedding, even when a query concerns only one semantic aspect, such as a class label or molecular property. Multiple heads can separate these a…

Text Retrieval

Effective LoRA Adapter Routing using Task Representations

2026-01-29 · Akash Dhasade, Anne-Marie Kermarrec, Igor Pavlovic, Diana Petrescu 외 arxiv

Low-rank adaptation (LoRA) enables parameter efficient specialization of large language models (LLMs) through modular adapters, resulting in rapidly growing public adapter pools spanning diverse tasks. Effectively using …

Representation over Routing: Diagnosing Temporal Routing Pathologies in Multi-Timescale PPO

2026-04-15 · Jing Sun arxiv

Temporal credit assignment in reinforcement learning is often approached by introducing value estimates at multiple discount factors. A natural next step is to let the actor dynamically route among these temporal heads, …

Reinforcement Learning