paper-with-me

홈 › Papers

Learning to Interpret Weight Differences in Language Models

2025-10-06 · Avichal Goel, Yoon Kim, Nir Shavit, Tony T. Wang arxiv

Finetuning (pretrained) language models is a standard approach for updating their internal parametric knowledge and specializing them to new tasks and domains. However, the corresponding model weight changes ("weight diffs") are not generally interpretable. While inspecting the finetuning dataset can give a sense of how the model might have changed, these datasets are often not publicly available or are too large to work with directly. Towards the goal of comprehensively understanding weight diffs in natural language, we introduce Diff Interpretation Tuning (DIT), a method that trains models to describe their own finetuning-induced modifications. Our approach uses synthetic, labeled weight diffs to train a DIT-adapter, which can be applied to a compatible finetuned model to make it describe how it has changed. We demonstrate in two proof-of-concept settings (reporting hidden behaviors and summarizing finetuned knowledge) that our method enables models to describe their finetuning-induced modifications using accurate natural language descriptions.

📄 PDF Abstract BibTeX arXiv:2510.05092

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Minding the Politeness Gap in Cross-cultural Communication

2025-06-18 · Yuka Machino, Matthias Hofer, Max Siegel, Joshua B. Tenenbaum 외

Misunderstandings in cross-cultural communication often arise from subtle differences in interpretation, but it is unclear whether these differences arise from the literal meanings assigned to words or from more general …

Conditional Hypothesis Generation for LLM-Based Text Analysis with Researcher-Specified Covariates

2026-06-02 · Paiheng Xu, Jing Liu, Wei Ai arxiv

A core goal of computational social science is to discover interpretable differences in how language varies across outcomes of interest, such as political affiliation or instructional quality. Recent LLM-based hypothesis…

AIPsychoBench: Understanding the Psychometric Differences between LLMs and Humans

2025-09-20 · Wei Xie, Shuoyoucheng Ma, Zhenhua Wang, Enze Wang 외 arxiv

Large Language Models (LLMs) with hundreds of billions of parameters have exhibited human-like intelligence by learning from vast amounts of internet-scale data. However, the uninterpretability of large-scale neural netw…

The Deleuzian Representation Hypothesis

2025-12-17 · Clément Cornet, Romaric Besançon, Hervé Le Borgne arxiv

We propose an alternative to sparse autoencoders (SAEs) as a simple and effective unsupervised method for extracting interpretable concepts from neural networks. The core idea is to cluster differences in activations, wh…

Interpretable Multi-dataset Evaluation for Named Entity Recognition

2020-11-13 · EMNLP 2020 11 · Jinlan Fu, PengFei Liu, Graham Neubig

With the proliferation of models for natural language processing tasks, it is even harder to understand the differences between models and their relative merits. Simply looking at differences between holistic metrics suc…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER