paper-with-me

홈 › Papers

CART: Context-Anchored Recurrent Transformer -- A Parameter-Efficient Architecture with Learned Stability

2026-05-31 · Chad A. Capps arxiv

We present CART (Context-Anchored Recurrent Transformer), a parameter-efficient language model that reuses a single shared core block R times across depth. Unlike prior looped transformers that recompute key-value tensors at every iteration, CART computes K and V once from a multi-layer prelude and has the recurrent core cross-attend to those frozen tensors via multi-head latent attention. A learned Linear Time-Invariant (LTI) gate keeps the recurrence stable: its spectral radius settles in a narrow band (rho in [0.79, 0.83]) across all 36 fully-trained configurations. We evaluate CART on single consumer GPUs in two stages: a 64-configuration screen at 3,000 steps, then 36 configurations (P=6, R in {6,8,10}, three seeds) trained for 30,500 steps (~1B tokens). Two patterns hold across widths d in {256,512,768,1024}: prelude depth P dominates loop count R, and the Stage-1 ranking of R reverses at full training (R=6 becomes best at d>=512). At the binding d=1024 parameter-parity test, CART does not beat a parameter-matched dense baseline, losing by 1-2% at stored-parameter parity and by ~10% at effective-parameter parity. Diagnostic ablations split the effective-parameter gap into ~5% from weight sharing and a residual ~5% from the heterogeneous prelude/anchor/core/coda framing; the recurrent-core machinery (hyper-connections, LTI gate, loop-index embedding) is individually vestigial. Variable-R inference degrades on both sides of the trained R, a negative result for test-time depth scaling under this recipe.

📄 PDF Abstract BibTeX arXiv:2606.01495

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CARTE: A Benchmark for Mapping Language Model Knowledge Across France

2026-06-01 · Sarah Almeida Carneiro, Christos Xypolopoulos, Xiao Fei, Yang Zhang 외 arxiv

We introduce CARTE 1 (Culturally Anchored Regional-Territorial Evaluation), a multiplechoice benchmark for evaluating the ability of large language models (LLMs) to perform fine-grained reasoning over geographically grou…

Cartoondiff: Training-free Cartoon Image Generation with Diffusion Transformer Models

2023-09-15 · Feihong He, Gang Li, Lingyu Si, Leilei Yan 외

Image cartoonization has attracted significant interest in the field of image generation. However, most of the existing image cartoonization techniques require re-training models using images of cartoon style. In this pa…

DenoisingImage Generation

Grocery to General Merchandise: A Cross-Pollination Recommender using LLMs and Real-Time Cart Context

2025-09-02 · Akshay Kekuda, Murali Mohana Krishna Dandu, Rimita Lahiri, Shiqin Cai 외 arxiv

Modern e-commerce platforms strive to enhance customer experience by providing timely and contextually relevant recommendations. However, recommending general merchandise to customers focused on grocery shopping -- such …

Mortality rate forecasting: can recurrent neural networks beat the Lee-Carter model?

2019-09-12 · Gábor Petneházi, József Gáll

This article applies a long short-term memory recurrent neural network to mortality rate forecasting. The model can be trained jointly on the mortality rate history of different countries, ages, and sexes. The RNN-based …

Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior

2026-05-26 · Zeyi Huang, Xuehai He, LiLiang Ren, Yiping Wang 외 arxiv

We study Latent Recurrent Transformer (LRT), a lightweight augmentation of autoregressive transformers that reuses a high-level source-layer hidden state from the previous token as recurrent memory for the next token. Be…