Multi-way VNMT for UGC: Improving Robustness and Capacity via Mixture Density Networks
This work presents a novel Variational Neural Machine Translation (VNMT) architecture with enhanced robustness properties, which we investigate through a detailed case-study addressing noisy French user-generated content (UGC) translation to English. We show that the proposed model, with results comparable or superior to state-of-the-art VNMT, improves performance over UGC translation in a zero-shot evaluation scenario while keeping optimal translation scores on in-domain test sets. We elaborate on such results by visualizing and explaining how neural learning representations behave when processing UGC noise. In addition, we show that VNMT enforces robustness to the learned embeddings, which can be later used for robust transfer learning approaches.
Code (0)
등록된 구현이 없습니다.
Tasks
Machine TranslationTransfer LearningTranslationSimilar Papers 제목 키워드 기반
Variational Neural Machine Translation with Normalizing Flows
Variational Neural Machine Translation (VNMT) is an attractive framework for modeling the generation of target translations, conditioned not only on the source sentence but also on some latent random variables. The laten…
Machine TranslationNMTSentenceTranslationMoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition
Large language models (LLMs) have recently shown strong potential in audio-visual speech recognition (AVSR), but their high computational demands and sensitivity to token granularity limit their practicality in resource-…
Audio-Visual Speech RecognitionRepresentation LearningLinear pretraining in recurrent mixture density networks
We present a method for pretraining a recurrent mixture density network (RMDN). We also propose a slight modification to the architecture of the RMDN-GARCH proposed by Nikolaev et al. [2012]. The pretraining method helps…
Generating Multiple Hypotheses for 3D Human Pose Estimation with Mixture Density Network
3D human pose estimation from a monocular image or 2D joints is an ill-posed problem because of depth ambiguity and occluded joints. We argue that 3D human pose estimation from a monocular input is an inverse problem whe…
3D Human Pose EstimationMonocular 3D Human Pose EstimationMulti-Hypotheses 3D Human Pose EstimationPose EstimationMulti-Task Mixture Density Graph Neural Networks for Predicting Cu-based Single-Atom Alloy Catalysts for CO2 Reduction Reaction
Graph neural networks (GNNs) have drawn more and more attention from material scientists and demonstrated a high capacity to establish connections between the structure and properties. However, with only unrelaxed struct…