paper-with-me

홈 › Papers

Beyond the Black Box: Identifiable Interpretation and Control in Generative Models via Causal Minimality

2025-12-11 · Lingjing Kong, Shaoan Xie, Guangyi Chen, Yuewen Sun, Xiangchen Song, Eric P. Xing, Kun Zhang arxiv

Deep generative models, while revolutionizing fields like image and text generation, largely operate as opaque ``black boxes'', hindering human understanding, control, and alignment. While methods like sparse autoencoders (SAEs) show remarkable empirical success, they often lack theoretical guarantees, risking subjective insights. Our primary objective is to establish a principled foundation for interpretable generative models. We demonstrate that the principle of causal minimality -- favoring the simplest causal explanation -- can endow the latent representations of modern generative models with clear causal interpretation and robust, component-wise identifiable control. We introduce a novel theoretical framework for hierarchical selection models, where higher-level concepts emerge from the constrained composition of lower-level variables, better capturing the complex dependencies in data generation. Under theoretically derived minimality conditions, we show that learned representations can be equivalent to the true latent variables of the data-generating process. Empirically, applying these constraints to leading text-to-image diffusion models allows us to extract their innate hierarchical concept graphs, offering fresh insights into their internal knowledge organization. Furthermore, these causally grounded concepts serve as levers for fine-grained model steering, paving the way for transparent, reliable systems.

📄 PDF Abstract BibTeX arXiv:2512.10720

Code (0)

등록된 구현이 없습니다.

Tasks

Text Generation

Similar Papers 제목 키워드 기반

Global Model Interpretation via Recursive Partitioning

2018-02-11 · Chengliang Yang, Anand Rangarajan, Sanjay Ranka

In this work, we propose a simple but effective method to interpret black-box machine learning models globally. That is, we use a compact binary tree, the interpretation tree, to explicitly represent the most important d…

BIG-bench Machine Learningmodel

On the Non-Identifiability of Steering Vectors in Large Language Models

2026-02-06 · Sohan Venkatesh, Ashish Mahendran Kurapath arxiv

Activation steering methods are widely used to control large language model (LLM) behavior and are often interpreted as revealing meaningful internal representations. This interpretation assumes that steering directions …

Opening the Black Box of Local Projections

2025-05-18 · Philippe Goulet Coulombe, Karin Klieber

Local projections (LPs) are widely used in empirical macroeconomics to estimate impulse responses to policy interventions. Yet, in many ways, they are black boxes. It is often unclear what mechanism or historical episode…

Kernel-Gradient Drifting Models

2026-05-11 · Maria Esteban-Casadevall, Jorge Carrasco-Pollo, Max Welling, Jan-Willem van de Meent 외 arxiv

We propose kernel-gradient drifting, a one-step generative modeling framework that replaces the fixed Euclidean displacement direction in drifting models with directions induced by the kernel itself. Standard drifting is…

Hijack-GAN: Unintended-Use of Pretrained, Black-Box GANs

2020-11-28 · CVPR 2021 1 · Hui-Po Wang, Ning Yu, Mario Fritz

While Generative Adversarial Networks (GANs) show increasing performance and the level of realism is becoming indistinguishable from natural images, this also comes with high demands on data and computation. We show that…

Image GenerationUnconditional Image Generation