paper-with-me

Papers

How deep is deep enough? -- Quantifying class separability in the hidden layers of deep neural networks

2018-11-05 · Achim Schilling, Claus Metzner, Jonas Rietsch, Richard Gerum, Holger Schulze, Patrick Krauss

Deep neural networks typically outperform more traditional machine learning models in their ability to classify complex data, and yet is not clear how the individual hidden layers of a deep network contribute to the overall classification performance. We thus introduce a Generalized Discrimination Value (GDV) that measures, in a non-invasive manner, how well different data classes separate in each given network layer. The GDV can be used for the automatic tuning of hyper-parameters, such as the width profile and the total depth of a network. Moreover, the layer-dependent GDV(L) provides new insights into the data transformations that self-organize during training: In the case of multi-layer perceptrons trained with error backpropagation, we find that classification of highly complex data sets requires a temporal {\em reduction} of class separability, marked by a characteristic 'energy barrier' in the initial part of the GDV(L) curve. Even more surprisingly, for a given data set, the GDV(L) is running through a fixed 'master curve', independently from the total number of network layers. Furthermore, applying the GDV to Deep Belief Networks reveals that also unsupervised training with the Contrastive Divergence method can systematically increase class separability over tens of layers, even though the system does not 'know' the desired class labels. These results indicate that the GDV may become a useful tool to open the black box of deep learning.

📄 PDF Abstract BibTeX arXiv:1811.01753

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Understanding Deep Neural Networks via Linear Separability of Hidden Layers

2023-07-26 · Chao Zhang, Xinyu Chen, Wensheng Li, Lixue Liu 외

In this paper, we measure the linear separability of hidden layer outputs to study the characteristics of deep neural networks. In particular, we first propose Minkowski difference based linear separability measures (MD-…

Unifying Attention Heads and Task Vectors via Hidden State Geometry in In-Context Learning

2025-05-24 · Haolin Yang, Hakaze Cho, Yiqiao Zhong, Naoya Inoue

The unusual properties of in-context learning (ICL) have prompted investigations into the internal mechanisms of large language models. Prior work typically focuses on either special attention heads or task vectors at sp…

In-Context Learning

Hidden Classification Layers: Enhancing linear separability between classes in neural networks layers

2023-06-09 · Andrea Apicella, Francesco Isgrò, Roberto Prevete

In the context of classification problems, Deep Learning (DL) approaches represent state of art. Many DL approaches are based on variations of standard multi-layer feed-forward neural networks. These are also referred to…

image-classificationImage Classification

Separability is not the best goal for machine learning

2018-07-08 · Wlodzislaw Duch

Neural networks use their hidden layers to transform input data into linearly separable data clusters, with a linear or a perceptron type output layer making the final projection on the line perpendicular to the discrimi…

BIG-bench Machine Learning

Generalization of an Upper Bound on the Number of Nodes Needed to Achieve Linear Separability

2018-02-10 · Marjolein Troost, Katja Seeliger, Marcel van Gerven

An important issue in neural network research is how to choose the number of nodes and layers such as to solve a classification problem. We provide new intuitions based on earlier results by An et al. (2015) by deriving …