paper-with-me

홈 › Papers

As Good as New. How to Successfully Recycle English GPT-2 to Make Models for Other Languages

2020-12-10 · Findings (ACL) 2021 8 · Wietse de Vries, Malvina Nissim

Large generative language models have been very successful for English, but other languages lag behind, in part due to data and computational limitations. We propose a method that may overcome these problems by adapting existing pre-trained models to new languages. Specifically, we describe the adaptation of English GPT-2 to Italian and Dutch by retraining lexical embeddings without tuning the Transformer layers. As a result, we obtain lexical embeddings for Italian and Dutch that are aligned with the original English lexical embeddings. Additionally, we scale up complexity by transforming relearned lexical embeddings of GPT-2 small to the GPT-2 medium embedding space. This method minimises the amount of training and prevents losing information during adaptation that was learned by GPT-2. English GPT-2 models with relearned lexical embeddings can generate realistic sentences in Italian and Dutch. Though on average these sentences are still identifiable as artificial by humans, they are assessed on par with sentences generated by a GPT-2 model fully trained from scratch.

📄 PDF Abstract BibTeX arXiv:2012.05628

Code (1)

wietsedv/gpt2-recycle 공식 구현 tf

Methods 이 논문이 사용한 방법론

Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Position-Wise Feed-Forward Layer 설명 없음
Cosine Annealing Cosine Annealing is a type of learning rate schedule that has the effect of starting with a large learning rate that is relatively rapidly decreased to a minimum value before…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Adam 설명 없음
Residual Connection 설명 없음
Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…

Similar Papers 제목 키워드 기반

Geometric Semantic Genetic Programming Algorithm and Slump Prediction

2017-09-18 · Juncai Xu, Zhenzhong Shen, Qingwen Ren, Xin Xie 외

Research on the performance of recycled concrete as building material in the current world is an important subject. Given the complex composition of recycled concrete, conventional methods for forecasting slump scarcely …

Prediction

If You Can't Use Them, Recycle Them: Optimizing Merging at Scale Mitigates Performance Tradeoffs

2024-12-05 · Muhammad Khalifa, Yi-Chern Tan, Arash Ahmadian, Tom Hosking 외

Model merging has shown great promise at combining expert models, but the benefit of merging is unclear when merging ``generalist'' models trained on many tasks. We explore merging in the context of large (~100B) models,…

Code GenerationInstruction Following

Recycle deep features for better object detection

2016-07-18 · Wei Li, Matthias Breier, Dorit Merhof

Aiming at improving the performance of existing detection algorithms developed for different applications, we propose a region regression-based multi-stage class-agnostic detection pipeline, whereby the existing algorith…

Objectobject-detectionObject Detectionregression

DuTongChuan: Context-aware Translation Model for Simultaneous Interpreting

2019-07-30 · Hao Xiong, Ruiqing Zhang, Chuanqiang Zhang, Zhongjun He 외

In this paper, we present DuTongChuan, a novel context-aware translation model for simultaneous interpreting. This model allows to constantly read streaming text from the Automatic Speech Recognition (ASR) model and simu…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)modelspeech-recognition+2

ReCycle: Resilient Training of Large DNNs using Pipeline Adaptation

2024-05-22 · Swapnil Gandhi, Mark Zhao, Athinagoras Skiadopoulos, Christos Kozyrakis

Training large Deep Neural Network (DNN) models requires thousands of GPUs over the course of several days or weeks. At this scale, failures are frequent and can have a big impact on training throughput. Utilizing spare …

GPU