paper-with-me

홈 › Papers

MoRFI: Monotonic Sparse Autoencoder Feature Identification

2026-04-29 · Dimitris Dimakopoulos, Shay B. Cohen, Ioannis Konstas arxiv

Large language models (LLMs) acquire most of their factual knowledge during the pre-training stage, through next token prediction. Subsequent stages of post-training often introduce new facts outwith the parametric knowledge, giving rise to hallucinations. While it has been demonstrated that supervised fine-tuning (SFT) on new knowledge may exacerbate the problem, the underlying mechanisms are still poorly understood. We conduct a controlled fine-tuning experiment, focusing on closed-book QA, and identify latent directions causally implicated in this degradation. Specifically, we fine-tune Llama 3.1 8B, Gemma 2 9B and Mistral 7B v03 on seven controlled mixtures of a single QA dataset, controlling for the percentage of new knowledge and number of training epochs. By measuring performance on the test set, we validate that incrementally introducing new knowledge increases hallucinations, with the effect being more pronounced with prolonged training. We leverage pre-trained sparse autoencoders (SAEs) to analyze residual stream activations across various checkpoints for each model and propose Monotonic Relationship Feature Identification (MoRFI) for capturing causally relevant latents. MoRFI filters SAE features that respond monotonically to controlled fine-tuning data mixtures of a target property. Our findings are consistent with exposure to unknown facts disrupting the model's ability to retrieve stored knowledge along a set of directions in the residual stream. Our pipeline reliably discovers them across distinct models, partially recovering lost knowledge through single-latent interventions.

📄 PDF Abstract BibTeX arXiv:2604.26866

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Evaluating Sparse Autoencoders: From Shallow Design to Matching Pursuit

2025-06-05 · Valérie Costa, Thomas Fel, Ekdeep Singh Lubana, Bahareh Tolooshams 외

Sparse autoencoders (SAEs) have recently become central tools for interpretability, leveraging dictionary learning principles to extract sparse, interpretable features from neural representations whose underlying structu…

Dictionary Learning

Omorfi --- Free and open source morphological lexical database for Finnish

2015-05-01 · WS 2015 5 · Tommi A. Pirinen
Machine TranslationMorphological Analysis

MorFiC: Fixing Value Miscalibration for Zero-Shot Quadruped Transfer

2026-03-15 · Prakhar Mishra, Amir Hossain Raj, Xuesu Xiao, Dinesh Manocha arxiv

Generalizing learned locomotion policies across quadrupedal robots with different morphologies remains a challenge. Policies trained on a single robot often fail when deployed on embodiments with different mass distribut…

Reinforcement Learning

A Robust SINDy Autoencoder for Noisy Dynamical System Identification

2026-04-06 · Kairui Ding arxiv

Sparse identification of nonlinear dynamics (SINDy) has been widely used to discover the governing equations of a dynamical system from data. It uses sparse regression techniques to identify parsimonious models of unknow…

Sparse Autoencoders Enable Scalable and Reliable Circuit Identification in Language Models

2024-05-21 · Charles O'Neill, Thang Bui

This paper introduces an efficient and robust method for discovering interpretable circuits in large language models using discrete sparse autoencoders. Our approach addresses key limitations of existing techniques, name…