paper-with-me

홈 › Papers

SVD-Softmax: Fast Softmax Approximation on Large Vocabulary Neural Networks

2017-12-01 · NeurIPS 2017 12 · Kyuhong Shim, Minjae Lee, Iksoo Choi, Yoonho Boo, Wonyong Sung

We propose a fast approximation method of a softmax function with a very large vocabulary using singular value decomposition (SVD). SVD-softmax targets fast and accurate probability estimation of the topmost probable words during inference of neural network language models. The proposed method transforms the weight matrix used in the calculation of the output vector by using SVD. The approximate probability of each word can be estimated with only a small part of the weight matrix by using a few large singular values and the corresponding elements for most of the words. We applied the technique to language modeling and neural machine translation and present a guideline for good approximation. The algorithm requires only approximately 20\% of arithmetic operations for an 800K vocabulary case and shows more than a three-fold speedup on a GPU.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

GPULanguage ModelingLanguage ModellingMachine TranslationTranslation

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…

Similar Papers 제목 키워드 기반

Navigating with Graph Representations for Fast and Scalable Decoding of Neural Language Models

2018-06-11 · NeurIPS 2018 12 · Minjia Zhang, Xiaodong Liu, Wenhan Wang, Jianfeng Gao 외

Neural language models (NLMs) have recently gained a renewed interest by achieving state-of-the-art performance across many natural language processing (NLP) tasks. However, NLMs are very computationally demanding largel…

DecoderLanguage ModelingLanguage ModellingMachine Translation+1

Learning to Screen for Fast Softmax Inference on Large Vocabulary Neural Networks

2018-10-29 · ICLR 2019 5 · Patrick H. Chen, Si Si, Sanjiv Kumar, Yang Li 외

Neural language models have been widely used in various NLP tasks, including machine translation, next word prediction and conversational agents. However, it is challenging to deploy these models on mobile devices due to…

ClusteringMachine TranslationPredictionTranslation

Real-time Neural-based Input Method

2018-10-19 · ICLR 2019 5 · Jiali Yao, Raphael Shu, Xinjian Li, Katsutoshi Ohtsuki 외

The input method is an essential service on every mobile and desktop devices that provides text suggestions. It converts sequential keyboard inputs to the characters in its target language, which is indispensable for Jap…

CPULanguage ModelingLanguage Modelling

Efficient softmax approximation for GPUs

2016-09-14 · ICML 2017 8 · Edouard Grave, Armand Joulin, Moustapha Cissé, David Grangier 외

We propose an approximate strategy to efficiently train neural network based language models over very large vocabularies. Our approach, called adaptive softmax, circumvents the linear dependency on the vocabulary size b…

GroupReduce: Block-Wise Low-Rank Approximation for Neural Language Model Shrinking

2018-06-18 · NeurIPS 2018 12 · Patrick H. Chen, Si Si, Yang Li, Ciprian Chelba 외

Model compression is essential for serving large deep neural nets on devices with limited resources or applications that require real-time responses. As a case study, a state-of-the-art neural language model usually cons…

Language ModelingLanguage ModellingModel CompressionQuantization