paper-with-me

Papers

Engineering Verifiable Modularity in Transformers via Per-Layer Supervision

2026-03-08 · J. Clayton Kerce arxiv

Transformers resist surgical control. Ablating an attention head identified as critical for capitalization produces minimal behavioral change because distributed redundancy compensates for damage. This Hydra effect renders interpretability illusory: we may identify components through correlation, but cannot predict or control their causal role. We demonstrate that architectural interventions can expose hidden modularity. Our approach combines dual-stream processing separating token and contextual representations, per-layer supervision providing independent gradient signal at each depth, and gated attention regularizing toward discrete activation patterns. When trained with per-layer supervision, models produce ablation effects 5 to 23 times larger than architecturally identical controls trained with standard objectives. This enables 4 times greater control leverage on targeted behaviors: scaling identified attention heads produces smooth, predictable changes in model output. The key finding is architectural. Without per-layer supervision, ablation damage concentrates near zero with low variance (Winograd standard deviation 0.63%). With per-layer supervision, effects spread widely (standard deviation 6.32%), revealing which predictions depend on which circuits. The larger variance is not measurement noise but the signature of unmasked modularity. We validate our approach through three components: engineered features that capture computational dynamics rather than vocabulary structure (validated by near-zero correlation with raw activation clustering), an architecture providing positive control for modularity, and causal experiments demonstrating functional reorganization where different tasks route through different attention heads. This es tablishes a methodology for transforming interpretability from passive observation to active control.

📄 PDF Abstract BibTeX arXiv:2603.18029

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently

2025-11-22 · Bochen Lyu, Yiyang Jia, Xiaohao Cai, Zhanxing Zhu arxiv

Transformers can acquire Chain-of-Thought (CoT) capabilities to solve reasoning tasks via fine-tuning. Reinforcement learning (RL) and supervised fine-tuning (SFT) are two primary approaches to this end. In this work, we…

Reinforcement Learning

EngTrace: A Symbolic Benchmark for Verifiable Process Supervision of Engineering Reasoning

2025-11-03 · Ayesha Gull, Muhammad Usman Safder, Rania Elbadry, Fan Zhang 외 arxiv

Large Language Models (LLMs) are increasingly entering specialized, safety-critical engineering workflows governed by strict quantitative standards and immutable physical laws, making rigorous evaluation of their reasoni…

Stream separation improves Bregman conditioning in transformers

2026-03-22 · James Clayton Kerce arxiv

Linear methods for steering transformer representations, including probing, activation engineering, and concept erasure, implicitly assume the geometry of representation space is Euclidean. Park et al. [Park et al., 2026…

Understanding the Dynamics of DNNs Using Graph Modularity

2021-11-24 · Yao Lu, Wen Yang, Yunzhe Zhang, Zuohui Chen 외

There are good arguments to support the claim that deep neural networks (DNNs) capture better feature representations than the previous hand-crafted feature engineering, which leads to a significant performance improveme…

Feature Engineering

Emergent Modularity in Pre-trained Transformers

2023-05-28 · Zhengyan Zhang, Zhiyuan Zeng, Yankai Lin, Chaojun Xiao 외

This work examines the presence of modularity in pre-trained Transformers, a feature commonly found in human brains and thought to be vital for general intelligence. In analogy to human brains, we consider two main chara…

Mixture-of-Experts