paper-with-me

홈 › Papers

The Context-Ready Transformer

2026-06-25 · Mahesh Godavarti arxiv

We introduce the context-ready transformer, a new recurrent neural network architecture built from a D-layer transformer block that pre-contextualizes each token before it enters the block. During left-to-right generation, a correction network combines the previous position's block output -- a cached summary of past context -- with the current token embedding, so the tokenenters the block already contextualized rather than as a raw embedding. At sequential inference, the correction chain makes the architecture a recurrent neural network. For training, we unroll the correction process K times over the full sequence, processing all positions in parallel at each step. A pretrained transformer can also be converted to a context-ready model by adding a zero-initialized correction FFN and fine-tuning. We evaluate across widths, depths, block sizes, and two datasets, with all comparisons against standard transformers, variants, and ablations. A D=5 model beats a 12-layer transformer while generating 1.7x faster on an A100. With K=10, a single-layermodel (D=1) beats a 6-layer transformer with a 2.6x inference speedup, and sequential inference matches parallel K=10 to within 0.01 PPL. The architecture benefits most from wide representations and long contexts. On a pointer-chasing task, D=1 trained with BPTT solves all 10 composition levels, while standard transformers exhibit staircase-like depth dependence.

📄 PDF Abstract BibTeX arXiv:2606.27538

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Segmenter: Transformer for Semantic Segmentation

2021-05-12 · ICCV 2021 10 · Robin Strudel, Ricardo Garcia, Ivan Laptev, Cordelia Schmid

Image segmentation is often ambiguous at the level of individual image patches and requires contextual information to reach label consensus. In this paper we introduce Segmenter, a transformer model for semantic segmenta…

Decoderimage-classificationImage ClassificationImage Segmentation+3

CodeSSM: Towards State Space Models for Code Understanding

2025-05-02 · Shweta Verma, Abhinav Anand, Mira Mezini

Although transformers are widely used for various code-specific tasks, they have some significant limitations. In this paper, we investigate State Space Models (SSMs) as a potential alternative to transformers for code u…

Clone DetectionLanguage ModelingLanguage ModellingMasked Language Modeling+3

Can Transformers Learn Full Bayesian Inference in Context?

2025-01-28 · Arik Reuter, Tim G. J. Rudner, Vincent Fortuin, David Rügamer

Transformers have emerged as the dominant architecture in the field of deep learning, with a broad range of applications and remarkable in-context learning (ICL) capabilities. While not yet fully understood, ICL has alre…

Bayesian InferenceIn-Context LearningVariational Inference

Algorithmic Capabilities of Random Transformers

2024-10-06 · Ziqian Zhong, Jacob Andreas

Trained transformer models have been found to implement interpretable procedures for tasks like arithmetic and associative recall, but little is understood about how the circuits that implement these procedures originate…

Text Generation

Attention Itself Could Retrieve.RetrieveVGGT: Training-Free Long Context Streaming 3D Reconstruction via Query-Key Similarity Retrieval

2026-05-10 · Zichen Zou, Xiaosong Jia, Zuxuan Wu, Yu-Gang Jiang arxiv

Visual Geometry Grounded Transformer (VGGT) advances 3D reconstruction via scalable Transformer architecture, but the quadratic complexity of global attention prevents long context application. StreamVGGT enables streami…

3D Reconstruction