paper-with-me

홈 › Papers

Causal Circuit Tracing Reveals Distinct Computational Architectures in Single-Cell Foundation Models: Inhibitory Dominance, Biological Coherence, and Cross-Model Convergence

2026-03-02 · Ihor Kendiukhov arxiv

Motivation: Sparse autoencoders (SAEs) decompose foundation model activations into interpretable features, but causal feature-to-feature interactions across network depth remain unknown for biological foundation models. Results: We introduce causal circuit tracing by ablating SAE features and measuring downstream responses, and apply it to Geneformer V2-316M and scGPT whole-human across four conditions (96,892 edges, 80,191 forward passes). Both models show approximately 53 percent biological coherence and 65 to 89 percent inhibitory dominance, invariant to architecture and cell type. scGPT produces stronger effects (mean absolute d = 1.40 vs. 1.05) with more balanced dynamics. Cross-model consensus yields 1,142 conserved domain pairs (10.6x enrichment, p < 0.001). Disease-associated domains are 3.59x more likely to be consensus. Gene-level CRISPRi validation shows 56.4 percent directional accuracy, confirming co-expression rather than causal encoding.

📄 PDF Abstract BibTeX arXiv:2603.01752

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Circuit Tracing in Vision-Language Models: Understanding the Internal Mechanisms of Multimodal Thinking

2026-02-23 · Jingcheng Yang, Tianhu Xiong, Shengyi Qian, Klara Nahrstedt 외 arxiv

Vision-language models (VLMs) are powerful but remain opaque black boxes. We introduce the first framework for transparent circuit tracing in VLMs to systematically analyze multimodal reasoning. By utilizing transcoders,…

Mathematical ReasoningMultimodal Reasoning

CircuitProbe: Tracing Visual Temporal Evidence Flow in Video Language Models

2025-07-25 · Yiming Zhang, Zhuokai Zhao, Chengzhang Yu, Kun Wang 외 arxiv

Autoregressive large vision--language models (LVLMs) interface video and language by projecting video features into the LLM's embedding space as continuous visual token embeddings. However, it remains unclear where tempo…

From Circuit Evidence to Mechanistic Theory: An Inductive Logic Approach

2026-05-20 · Nura Aljaafari, Danilo S. Carvalho, Andre Freitas arxiv

Mechanistic interpretability produces circuit-level causal analyses of neural network behaviour, but discovered circuits often remain isolated experimental artefacts: there is no shared formal representation for what cir…

Inductive logic programming

Exhaustive Circuit Mapping of a Single-Cell Foundation Model Reveals Massive Redundancy, Heavy-Tailed Hub Architecture, and Layer-Dependent Differentiation Control

2026-03-12 · Ihor Kendiukhov arxiv

Mechanistic interpretability of biological foundation models has relied on selective feature sampling, pairwise interaction testing, and observational trajectory analysis. Each of these can introduce systematic bias. Her…

The Quantum Sieve Tracer: A Hybrid Framework for Layer-Wise Activation Tracing in Large Language Models

2026-02-06 · Jonathan Pan arxiv

Mechanistic interpretability aims to reverse-engineer the internal computations of Large Language Models (LLMs), yet separating sparse semantic signals from high-dimensional polysemantic noise remains a significant chall…