paper-with-me

홈 › Papers

Optimizing transformer-based machine translation model for single GPU training: a hyperparameter ablation study

2023-08-11 · Luv Verma, Ketaki N. Kolhatkar

In machine translation tasks, the relationship between model complexity and performance is often presumed to be linear, driving an increase in the number of parameters and consequent demands for computational resources like multiple GPUs. To explore this assumption, this study systematically investigates the effects of hyperparameters through ablation on a sequence-to-sequence machine translation pipeline, utilizing a single NVIDIA A100 GPU. Contrary to expectations, our experiments reveal that combinations with the most parameters were not necessarily the most effective. This unexpected insight prompted a careful reduction in parameter sizes, uncovering "sweet spots" that enable training sophisticated models on a single GPU without compromising translation quality. The findings demonstrate an intricate relationship between hyperparameter selection, model size, and computational resource needs. The insights from this study contribute to the ongoing efforts to make machine translation more accessible and cost-effective, emphasizing the importance of precise hyperparameter tuning over mere scaling.

📄 PDF Abstract BibTeX arXiv:2308.06017

Code (0)

등록된 구현이 없습니다.

Tasks

GPUMachine TranslationTranslation

Similar Papers 제목 키워드 기반

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

2019-10-01 · WS 2019 11 · Kenton Murray, Jeffery Kinnison, Toan Q. Nguyen, Walter Scheirer 외

Neural sequence-to-sequence models, particularly the Transformer, are the state of the art in machine translation. Yet these neural networks are very sensitive to architecture and hyperparameter settings. Optimizing thes…

Machine TranslationTranslation

Optimizing Transformer for Low-Resource Neural Machine Translation

2020-11-04 · COLING 2020 8 · Ali Araabi, Christof Monz

Language pairs with limited amounts of parallel data, also known as low-resource languages, remain a challenge for neural machine translation. While the Transformer model has achieved significant improvements for many la…

Low Resource Neural Machine TranslationLow-Resource Neural Machine TranslationMachine TranslationTranslation

Maximum Proxy-Likelihood Estimation for Non-autoregressive Machine Translation

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Maximum Likelihood Estimation (MLE) is commonly used in machine translation, where models with higher likelihood are assumed to perform better in translation. However, this assumption does not hold in the non-autoregress…

Machine TranslationTranslation

Optimizing Deep Transformers for Chinese-Thai Low-Resource Translation

2022-12-24 · Wenjie Hao, Hongfei Xu, Lingling Mu, Hongying Zan

In this paper, we study the use of deep Transformer translation model for the CCMT 2022 Chinese-Thai low-resource machine translation task. We first explore the experiment settings (including the number of BPE merge oper…

Machine TranslationTranslation

Optimizing example selection for retrieval-augmented machine translation with translation memories

2024-05-23 · Maxime Bouthors, Josep Crego, François Yvon

Retrieval-augmented machine translation leverages examples from a translation memory by retrieving similar instances. These examples are used to condition the predictions of a neural decoder. We aim to improve the upstre…

DecoderMachine TranslationRetrievalSentence+1