paper-with-me

Papers

Applying Cyclical Learning Rate to Neural Machine Translation

2020-04-06 · Choon Meng Lee, Jianfeng Liu, Wei Peng

In training deep learning networks, the optimizer and related learning rate are often used without much thought or with minimal tuning, even though it is crucial in ensuring a fast convergence to a good quality minimum of the loss function that can also generalize well on the test dataset. Drawing inspiration from the successful application of cyclical learning rate policy for computer vision related convolutional networks and datasets, we explore how cyclical learning rate can be applied to train transformer-based neural networks for neural machine translation. From our carefully designed experiments, we show that the choice of optimizers and the associated cyclical learning rate policy can have a significant impact on the performance. In addition, we establish guidelines when applying cyclical learning rates to neural machine translation tasks. Thus with our work, we hope to raise awareness of the importance of selecting the right optimizers and the accompanying learning rate policy, at the same time, encourage further research into easy-to-use learning rate policies.

📄 PDF Abstract BibTeX arXiv:2004.02401

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationTranslation

Methods 이 논문이 사용한 방법론

Cyclical Learning Rate Policy A Cyclical Learning Rate Policy combines a linear learning rate decay with warm restarts. Image: ESPNetv2

Similar Papers 제목 키워드 기반

CycleDistill: Bootstrapping Machine Translation using LLMs with Cyclical Distillation

2025-06-24 · Deepon Halder, Thanmay Jayakumar, Raj Dabre

Large language models (LLMs), despite their ability to perform few-shot machine translation (MT), often lag behind dedicated MT systems trained on parallel corpora, which are crucial for high quality machine translation …

Machine TranslationTranslation

General Cyclical Training of Neural Networks

2022-02-17 · Leslie N. Smith

This paper describes the principle of "General Cyclical Training" in machine learning, where training starts and ends with "easy training" and the "hard training" happens during the middle epochs. We propose several mani…

Data AugmentationKnowledge Distillation

Applying SVGD to Bayesian Neural Networks for Cyclical Time-Series Prediction and Inference

2019-01-17 · Xinyu Hu, Paul Szerlip, Theofanis Karaletsos, Rohit Singh

A regression-based BNN model is proposed to predict spatiotemporal quantities like hourly rider demand with calibrated uncertainties. The main contributions of this paper are (i) A feed-forward deterministic neural netwo…

regressionSensitivityTime SeriesTime Series Analysis+1

Applying Automated Machine Translation to Educational Video Courses

2023-01-09 · Linden Wang

We studied the capability of automated machine translation in the online video education space by automatically translating Khan Academy videos with state-of-the-art translation models and applying text-to-speech synthes…

Machine TranslationSpeech Synthesistext-to-speechText to Speech+3

Applying Machine Translation Metrics to Student-Written Translations

2013-06-01 · WS 2013 6 · Lisa Michaud, Patricia Ann McCoy
Machine TranslationTranslation