paper-with-me

홈 › Papers

Self-organized Hierarchical Softmax

2017-07-26 · Yikang Shen, Shawn Tan, Chrisopher Pal, Aaron Courville

We propose a new self-organizing hierarchical softmax formulation for neural-network-based language models over large vocabularies. Instead of using a predefined hierarchical structure, our approach is capable of learning word clusters with clear syntactical and semantic meaning during the language model training process. We provide experiments on standard benchmarks for language modeling and sentence compression tasks. We find that this approach is as fast as other efficient softmax approximations, while achieving comparable or even better performance relative to similar full softmax models.

📄 PDF Abstract BibTeX arXiv:1707.08588

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingSentenceSentence Compression

Methods 이 논문이 사용한 방법론

Hierarchical Softmax 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…

Similar Papers 제목 키워드 기반

Intrinsic and Extrinsic Organized Attention: Softmax Invariance and Network Sparsity

2025-06-18 · Oluwadamilola Fasina, Ruben V. C. Pohle, Pei-Chun Su, Ronald R. Coifman

We examine the intrinsic (within the attention head) and extrinsic (amongst the attention heads) structure of the self-attention mechanism in transformers. Theoretical evidence for invariance of the self-attention mechan…

Global Hierarchical Neural Networks using Hierarchical Softmax

2023-08-02 · Jetze Schuurmans, Flavius Frasincar

This paper presents a framework in which hierarchical softmax is used to create a global hierarchical classifier. The approach is applicable for any classification task where there is a natural hierarchy among classes. W…

Classificationtext-classificationText Classification

Effectiveness of Hierarchical Softmax in Large Scale Classification Tasks

2018-12-13 · Abdul Arfat Mohammed, Venkatesh Umaashankar

Typically, Softmax is used in the final layer of a neural network to get a probability distribution for output classes. But the main problem with Softmax is that it is computationally expensive for large scale data sets …

ClassificationGeneral Classification

Hyperbolic Additive Margin Softmax with Hierarchical Information for Speaker Verification

2026-01-27 · Zhihua Fang, Liang He arxiv

Speaker embedding learning based on Euclidean space has achieved significant progress, but it is still insufficient in modeling hierarchical information within speaker features. Hyperbolic space, with its negative curvat…

Speaker Verification

Strategies for Training Large Vocabulary Neural Language Models

2015-12-15 · ACL 2016 8 · Welin Chen, David Grangier, Michael Auli

Training neural network language models over large vocabularies is still computationally very costly compared to count-based models such as Kneser-Ney. At the same time, neural language models are gaining popularity for …

Machine Translationspeech-recognitionSpeech RecognitionTranslation