paper-with-me

Papers

Transcoder-based Circuit Analysis for Interpretable Single-Cell Foundation Models

2025-09-18 · Sosuke Hosokawa, Toshiharu Kawakami, Satoshi Kodera, Masamichi Ito, Norihiko Takeda arxiv

Single-cell foundation models (scFMs) have demonstrated state-of-the-art performance on various tasks, such as cell-type annotation and perturbation response prediction, by learning gene regulatory networks from large-scale transcriptome data. However, a significant challenge remains: the decision-making processes of these models are less interpretable compared to traditional methods like differential gene expression analysis. Recently, transcoders have emerged as a promising approach for extracting interpretable decision circuits from large language models (LLMs). In this work, we train a transcoder on the cell2sentence (C2S) model, a state-of-the-art scFM. By leveraging the trained transcoder, we extract internal decision-making circuits from the C2S model. We demonstrate that the discovered circuits correspond to real-world biological mechanisms, confirming the potential of transcoders to uncover biologically plausible pathways within complex single-cell models.

📄 PDF Abstract BibTeX arXiv:2509.14723

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Transcoders Find Interpretable LLM Feature Circuits

2024-06-17 · Jacob Dunefsky, Philippe Chlenski, Neel Nanda

A key goal in mechanistic interpretability is circuit analysis: finding sparse subgraphs of models corresponding to specific behaviors or capabilities. However, MLP sublayers make fine-grained circuit analysis on transfo…

DifFRACT: Diffusion Feature Reconstruction and Attribution for Circuit Tracing

2026-06-14 · Artyom Mazur, Nina Konovalova, Aibek Alanov arxiv

Mechanistic interpretability seeks to explain neural network behavior by decomposing model computations into interpretable features and circuits. While transcoder-based circuit tracing has recently enabled detailed causa…

Image Generation

Transcoders Trace Visual Grounding and Hallucinations in Vision-Language Models

2026-05-21 · Dimitrios Damianos, Leon Voukoutis, Georgios Skyrianos, Vassilis Katsouros 외 arxiv

Generative Vision-Language Models (VLMs) perform well on multimodal reasoning, but how visual inputs are transformed to text remains poorly understood. Existing interpretability work on VLMs uses Sparse Autoencoders (SAE…

Multimodal ReasoningVisual Grounding

Transcoders for Investigating Deception in Language Models

2026-07-16 · Darius Lim, Nathan Leow, Xin Wei Chia arxiv

Transcoders have recently emerged as a promising approach for mechanistic interpretability (MI), enabling circuit-level analysis of model behaviour. In this paper, we investigate the use of transcoders to analyse decepti…

Protein Circuit Tracing via Cross-layer Transcoders

2026-02-12 · Darin Tsui, Kunal Talreja, Daniel Saeedi, Amirali Aghazadeh arxiv

Protein language models (pLMs) have emerged as powerful predictors of protein structure and function. However, the computational circuits underlying their predictions remain poorly understood. Recent mechanistic interpre…

Protein Design