paper-with-me

Subformer

2000년 도입 · 논문 3편에서 사용

Subformer is a Transformer that combines sandwich-style parameter sharing, which overcomes naive cross-layer parameter sharing in generative models, and self-attentive embedding factorization (SAFE). In SAFE, a small self-attention layer is used to reduce embedding parameter count.

출처: Subformer: Exploring Weight Sharing for Parameter Efficiency in Generative Transformers

소개 논문: Subformer: Exploring Weight Sharing for Parameter Efficiency in Generative Transformers

Transformers · Natural Language Processing