Isotropic Contextual Representations through Variational Regularization
Contextual language representations achieve state-of-the-art performance across various natural language processing tasks. However, these representations have been shown to suffer from the degeneration problem, i.e. they occupy a narrow cone in the latent space. This problem can be addressed by enforcing isotropy in the latent space. In analogy to variational autoencoders, we suggest applying a token-level variational loss to a Transformer architecture and introduce the prior distribution's standard deviation as model parameter to optimize isotropy. The encoder-decoder architecture allows for learning interpretable embeddings that can be decoded into text again. Extracted features at sentence-level achieve competitive results on benchmark classification tasks.
Code (0)
등록된 구현이 없습니다.
Tasks
DecoderSentenceMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Stable Anisotropic Regularization
Given the success of Large Language Models (LLMs), there has been considerable interest in studying the properties of model activations. The literature overwhelmingly agrees that LLM representations are dominated by a fe…
Space-adaptive anisotropic bivariate Laplacian regularization for image restoration
In this paper we present a new regularization term for variational image restoration which can be regarded as a space-variant anisotropic extension of the classical isotropic Total Variation (TV) regularizer. The propose…
Image RestorationLow Anisotropy Sense Retrofitting (LASeR) : Towards Isotropic and Sense Enriched Representations
Contextual word representation models have shown massive improvements on a multitude of NLP tasks, yet their word sense disambiguation capabilities remain poorly explained. To address this gap, we assess whether contextu…
Word Sense DisambiguationVariational Depth Superresolution Using Example-Based Edge Representations
In this paper we propose a novel method for depth image superresolution which combines recent advances in example based upsampling with variational superresolution based on a known blur kernel. Most traditional depth sup…
MIC: Maximizing Informational Capacity in Adaptive Representations via Isotropic Subspace Alignment
Although multi-scales representation learning enables elastic-dimension embeddings, nested subspaces often suffer from dimensional redundancy and spectral collapse. To address this, we introduce MIC, a framework that opt…
Representation Learning