Self-organized Hierarchical Softmax
We propose a new self-organizing hierarchical softmax formulation for neural-network-based language models over large vocabularies. Instead of using a predefined hierarchical structure, our approach is capable of learning word clusters with clear syntactical and semantic meaning during the language model training process. We provide experiments on standard benchmarks for language modeling and sentence compression tasks. We find that this approach is as fast as other efficient softmax approximations, while achieving comparable or even better performance relative to similar full softmax models.
Code (0)
등록된 구현이 없습니다.
Tasks
Language ModelingLanguage ModellingSentenceSentence CompressionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Intrinsic and Extrinsic Organized Attention: Softmax Invariance and Network Sparsity
We examine the intrinsic (within the attention head) and extrinsic (amongst the attention heads) structure of the self-attention mechanism in transformers. Theoretical evidence for invariance of the self-attention mechan…
Global Hierarchical Neural Networks using Hierarchical Softmax
This paper presents a framework in which hierarchical softmax is used to create a global hierarchical classifier. The approach is applicable for any classification task where there is a natural hierarchy among classes. W…
Classificationtext-classificationText ClassificationEffectiveness of Hierarchical Softmax in Large Scale Classification Tasks
Typically, Softmax is used in the final layer of a neural network to get a probability distribution for output classes. But the main problem with Softmax is that it is computationally expensive for large scale data sets …
ClassificationGeneral ClassificationHyperbolic Additive Margin Softmax with Hierarchical Information for Speaker Verification
Speaker embedding learning based on Euclidean space has achieved significant progress, but it is still insufficient in modeling hierarchical information within speaker features. Hyperbolic space, with its negative curvat…
Speaker VerificationStrategies for Training Large Vocabulary Neural Language Models
Training neural network language models over large vocabularies is still computationally very costly compared to count-based models such as Kneser-Ney. At the same time, neural language models are gaining popularity for …
Machine Translationspeech-recognitionSpeech RecognitionTranslation