paper-with-me

Papers

Subspace Match Probably Does Not Accurately Assess the Similarity of Learned Representations

2019-01-03 · Jeremiah Johnson

Learning informative representations of data is one of the primary goals of deep learning, but there is still little understanding as to what representations a neural network actually learns. To better understand this, subspace match was recently proposed as a method for assessing the similarity of the representations learned by neural networks. It has been shown that two networks with the same architecture trained from different initializations learn representations that at hidden layers show low similarity when assessed with subspace match, even when the output layers show high similarity and the networks largely exhibit similar performance on classification tasks. In this note, we present a simple example motivated by standard results in commutative algebra to illustrate how this can happen, and show that although the subspace match at a hidden layer may be 0, the representations learned may be isomorphic as vector spaces. This leads us to conclude that a subspace match comparison of learned representations may well be uninformative, and it points to the need for better methods of understanding learned representations.

📄 PDF Abstract BibTeX arXiv:1901.00884

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Compressed Computation is (probably) not Computation in Superposition

2026-06-12 · Jai Bhagat, Sara Molas-Medina, Giorgi Giglemiani, Stefan Heimersheim arxiv

We study whether the Compressed Computation (CC) toy model (Braun et al., 2025) is an instance of computation in superposition. The CC model appears to compute 100 ReLU functions with just 50 neurons, achieving a better …

Theory of matching pursuit

2008-12-01 · NeurIPS 2008 12 · Zakria Hussain, John S. Shawe-Taylor

We analyse matching pursuit for kernel principal components analysis by proving that the sparse subspace it produces is a sample compression scheme. We show that this bound is tighter than the KPCA bound of Shawe-Taylor …

SAD-LoRA: Spectral Alignment for Low-Rank Knowledge Distillation

2026-07-05 · Omer Tariq, Syed Muhammad Raza, Jeongbae Son arxiv

Distilling a fine-tuned teacher into a LoRA-adapted student is a standard recipe for parameter-efficient compression, but output-level KD does not explicitly control which rank-$r$ weight subspace the adapter occupies. W…

Knowledge Distillation

Generative Modeling of Discrete Data Using Geometric Latent Subspaces

2026-01-29 · Daniel Gonzalez-Alvarado, Jonas Cassel, Stefania Petra, Christoph Schnörr arxiv

We propose a geometric latent-subspace framework for generative modeling of discrete data. Specifically, we introduce latent subspaces in the exponential parameter space of product manifolds of categorical distributions …

Dimensionality Reduction

Probably Correct Optimal Stable Matching for Two-Sided Markets Under Uncertainty

2025-01-06 · Andreas Athanasopoulos, Anne-Marie George, Christos Dimitrakakis

We consider a learning problem for the stable marriage model under unknown preferences for the left side of the market. We focus on the centralized case, where at each time step, an online platform matches the agents, an…