Low-loss connection of weight vectors: distribution-based approaches
Recent research shows that sublevel sets of the loss surfaces of overparameterized networks are connected, exactly or approximately. We describe and compare experimentally a panel of methods used to connect two low-loss points by a low-loss curve on this surface. Our methods vary in accuracy and complexity. Most of our methods are based on "macroscopic" distributional assumptions, and some are insensitive to the detailed properties of the points to be connected. Some methods require a prior training of a "global connection model" which can then be applied to any pair of points. The accuracy of the method generally correlates with its complexity and sensitivity to the endpoint detail.
Code (1)
Tasks
SensitivitySimilar Papers 제목 키워드 기반
Unsupervised Learning of Distributional Relation Vectors
Word embedding models such as GloVe rely on co-occurrence statistics to learn vector representations of word meaning. While we may similarly expect that co-occurrence statistics can be used to capture rich information ab…
RelationRelation ExtractionWord EmbeddingsModeling Semantic Relatedness using Global Relation Vectors
Word embedding models such as GloVe rely on co-occurrence statistics from a large corpus to learn vector representations of word meaning. These vectors have proven to capture surprisingly fine-grained semantic and syntac…
RelationA Thesaurus for Biblical Hebrew
We built a thesaurus for Biblical Hebrew, with connections between roots based on phonetic, semantic, and distributional similarity. To this end, we apply established algorithms to find connections between headwords base…
Semantic SimilaritySemantic Textual SimilarityDropout Training for Support Vector Machines
Dropout and other feature noising schemes have shown promising results in controlling over-fitting by artificially corrupting the training data. Though extensive theoretical and empirical studies have been performed for …
Data AugmentationOn the Connection Between Learning Two-Layers Neural Networks and Tensor Decomposition
We establish connections between the problem of learning a two-layer neural network and tensor decomposition. We consider a model with feature vectors $\boldsymbol x \in \mathbb R^d$, $r$ hidden units with weights $\{\bo…
Tensor Decomposition