paper-with-me

홈 › Papers

The Dual-Stream Transformer: Channelized Architecture for Interpretable Language Modeling

2026-03-08 · J. Clayton Kerce, Alexis Fox arxiv

Standard transformers entangle all computation in a single residual stream, obscuring which components perform which functions. We introduce the Dual-Stream Transformer, which decomposes the residual stream into two functionally distinct components: a token stream updated by attention and a context stream updated by feed-forward networks. Information flow between attention heads is controlled through a hierarchy of mixing strategies, from fully independent (maximum interpretability) to dense (standard transformer behavior). This design exposes a tunable tradeoff between interpretability and performance. We measure this tradeoff on language modeling tasks at 29M parameters. Fully independent head mixing increases validation loss by 8\% relative to dense baselines. The recommended Kronecker mixing strategy, which permits scalar communication between heads while preserving within-head structure, costs only 2.5\%. All configurations maintain functional generation under attention amplification (scaling logits by factors up to 16 at inference time), with degradation ranging from 16\% to 27\%. This robustness suggests the architectures learn discrete algorithms that operate independently of soft probabilistic mixing. The architecture provides a foundation for interpretable language models where internal structure is exposed by design. \footnote{This work was partially supported by DARPA Contract HR001125C0302.}

📄 PDF Abstract BibTeX arXiv:2603.07461

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A deep-learning-based surrogate model for data assimilation in dynamic subsurface flow problems

2019-08-16 · Meng Tang, Yimin Liu, Louis J. Durlofsky

A deep-learning-based surrogate model is developed and applied for predicting dynamic subsurface flow in channelized geological models. The surrogate model is based on deep convolutional and recurrent neural network arch…

A Channelized Binning Method for Extraction of Dominant Color Pixel Value

2016-05-28 · Siddu P Algur, N H Ayachit, Vivek R

The Color is one of the most important and easily identifiable features for describing the visual content. The MPEG standard has developed a number of descriptors that covers different aspects of the visual content. The …

Quantization

Closing the Loop: PID Feedback Control for Interpretable Activation Steering in Symbolic Music Generation

2026-06-17 · Ioannis Prokopiou, Pantelis Vikatos, Maximos Kaliakatsos-Papakostas, Theodoros Giannakopoulos 외 arxiv

Transformer-based architectures have significantly advanced the generation of complex symbolic sequences, yet a significant gap remains in achieving fine-grained, interpretable control over discrete signal attributes. Th…

Music Generation

Latent Space Disentanglement via Activation Steering for Interpretable Attribute Control in Symbolic Music Generation

2026-05-29 · Ioannis Prokopiou, Pantelis Vikatos, Maximos Kaliakatsos-Papakostas, Theodoros Giannakopoulos 외 arxiv

Transformer-based architectures have significantly advanced the generation of complex symbolic sequences, yet a significant gap remains in achieving fine-grained, interpretable control over discrete signal attributes. Th…

Music Generation

Integration of spatio-temporal contrast sensitivity with a multi-slice channelized Hotelling observer

2013-04-04 · Ali N. Avanaki, Kathryn S. Espig, Cedric Marchessoux, Elizabeth A. Krupinski 외

Barten's model of spatio-temporal contrast sensitivity function of human visual system is embedded in a multi-slice channelized Hotelling observer. This is done by 3D filtering of the stack of images with the spatio-temp…

Sensitivity