Embedding Dropout
2000년 도입 · 논문 64편에서 사용
Embedding Dropout is equivalent to performing dropout on the embedding matrix at a word level, where the dropout is broadcast across all the word vector’s embedding. The remaining non-dropped-out word embeddings are scaled by $\frac{1}{1-p\_{e}}$ where $p\_{e}$ is the probability of embedding dropout. As the dropout occurs on the embedding matrix that is used for a full forward and backward pass, this means that all occurrences of a specific word will disappear within that pass, equivalent to performing variational dropout on the connection between the one-hot embedding and the embedding lookup. Source: Merity et al, Regularizing and Optimizing LSTM Language Models
출처: A Theoretically Grounded Application of Dropout in Recurrent Neural Networks
소개 논문: A Theoretically Grounded Application of Dropout in Recurrent Neural Networks
Regularization · General