paper-with-me

Papers

VFUSE: Virulent Feature Understanding with Sparse autoEncoders

2026-06-08 · Michael Yu, Matthew L. Olson arxiv

Generative models have shown remarkable progress in a variety of domains such as protein design, but such power enables the opaque generation of hazardous proteins. In this work, we introduce VFUSE (Virulent Feature Understanding with Sparse autoEncoders), a mechanistic interpretability approach that trains SAEs on diffusion-transformer activations to audit protein models for hazard-aware features. We apply VFUSE to RoseTTAFold3 and RFDiffusion3, popular open-weight models for protein folding and synthesis. We find that for certain blocks, linear probes detect hazardous designs significantly better when fit in the SAE latent space over the original model's representations: improving interpretability without sacrificing model performance. Furthermore, we identify monosemantic features from the SAE that fire only on hazardous designs at up to AUROC $0.84$ ($q < 10^{-13}$). To our knowledge this is the first SAE trained on an all-atom diffusion model and the first feature-level virulence audit of a protein design model, paving the way towards safe and interpretable protein design.

📄 PDF Abstract BibTeX arXiv:2606.10080

Code (0)

등록된 구현이 없습니다.

Tasks

Protein Design

Similar Papers 제목 키워드 기반

vFusedSeg3D: 3rd Place Solution for 2024 Waymo Open Dataset Challenge in Semantic Segmentation

2024-08-09 · Osama Amjad, Ammad Nadeem

In this technical study, we introduce VFusedSeg3D, an innovative multi-modal fusion system created by the VisionRD team that combines camera and LiDAR data to significantly enhance the accuracy of 3D perception. VFusedSe…

3D Semantic SegmentationSemantic Segmentation

MVFuseNet: Improving End-to-End Object Detection and Motion Forecasting through Multi-View Fusion of LiDAR Data

2021-04-21 · Ankit Laddha, Shivam Gautam, Stefan Palombo, Shreyash Pandey 외

In this work, we propose \textit{MVFuseNet}, a novel end-to-end method for joint object detection and motion forecasting from a temporal sequence of LiDAR data. Most existing methods operate in a single view by projectin…

Motion Forecastingobject-detectionObject Detection

An Approach for Molecular and Biological Characterizations of Virulent Influenza A Viruses

2024-11-07 · Meitner Cadena, Alejandro Yerovi

We propose an approach based on a combination of physical, chemical, and mathematical methods to identify and characterize virulent influenza A viruses (IAVs) through the analysis of the hemagglutinin protein. These meth…

Position

AdaptiveK: Complexity-Driven Sparse Autoencoders for Interpretable Language Model Representations

2025-08-24 · Yifei Yao, Hanrong Zhang, Mengnan Du arxiv

Understanding the internal representations of large language models (LLMs) remains a central challenge for interpretability research. Sparse autoencoders (SAEs) offer a promising solution by decomposing activations into …

Understanding Refusal in Language Models with Sparse Autoencoders

2025-05-29 · Wei Jie Yeo, Nirmalendu Prakash, Clement Neo, Roy Ka-Wei Lee 외

Refusal is a key safety behavior in aligned language models, yet the internal mechanisms driving refusals remain opaque. In this work, we conduct a mechanistic study of refusal in instruction-tuned LLMs using sparse auto…