paper-with-me

홈 › Papers

Controllable Dual Skew Divergence Loss for Neural Machine Translation

2019-08-22 · Zuchao Li, Hai Zhao, Yingting Wu, Fengshun Xiao, Shu Jiang

In sequence prediction tasks like neural machine translation, training with cross-entropy loss often leads to models that overgeneralize and plunge into local optima. In this paper, we propose an extended loss function called \emph{dual skew divergence} (DSD) that integrates two symmetric terms on KL divergences with a balanced weight. We empirically discovered that such a balanced weight plays a crucial role in applying the proposed DSD loss into deep models. Thus we eventually develop a controllable DSD loss for general-purpose scenarios. Our experiments indicate that switching to the DSD loss after the convergence of ML training helps models escape local optima and stimulates stable performance improvements. Our evaluations on the WMT 2014 English-German and English-French translation tasks demonstrate that the proposed loss as a general and convenient mean for NMT training indeed brings performance improvement in comparison to strong baselines.

📄 PDF Abstract BibTeX arXiv:1908.08399

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationNMTTranslation

Similar Papers 제목 키워드 기반

Divergences induced by dual subtractive and divisive normalizations of exponential families and their convex deformations

2023-12-20 · Frank Nielsen

Exponential families are statistical models which are the workhorses in statistics, information theory, and machine learning among others. An exponential family can either be normalized subtractively by its cumulant or f…

On $w$-mixtures: Finite convex combinations of prescribed component distributions

2017-08-02 · Frank Nielsen, Richard Nock

We consider the space of $w$-mixtures which is defined as the set of finite statistical mixtures sharing the same prescribed component distributions closed under convex combinations. The information geometry induced by t…

$\alpha$-VAEs : Optimising variational inference by learning data-dependent divergence skew

2021-06-02 · ICML Workshop INNF 2021 7 · Jacob Deasy, Tom Andrew McIver, Nikola Simidjievski, Pietro Lio

The {\em skew-geometric Jensen-Shannon divergence} $\left(\textrm{JS}^{\textrm{G}_{\alpha}}\right)$ allows for an intuitive interpolation between forward and reverse Kullback-Leibler (KL) divergence based on the skew pa…

DenoisingVariational Inference

Generalized Bregman and Jensen divergences which include some f-divergences

2018-08-19 · Tomohiro Nishiyama

In this paper, we introduce new classes of divergences by extending the definitions of the Bregman divergence and the skew Jensen divergence. These new divergence classes (g-Bregman divergence and skew g-Jensen divergenc…

$α$-Geodesical Skew Divergence

2021-03-31 · Masanari Kimura, Hideitsu Hino

The asymmetric skew divergence smooths one of the distributions by mixing it, to a degree determined by the parameter $\lambda$, with the other distribution. Such divergence is an approximation of the KL divergence that …