paper-with-me

홈 › Papers

Guiding Teacher Forcing with Seer Forcing for Neural Machine Translation

2021-06-12 · ACL 2021 5 · Yang Feng, Shuhao Gu, Dengji Guo, Zhengxin Yang, Chenze Shao

Although teacher forcing has become the main training paradigm for neural machine translation, it usually makes predictions only conditioned on past information, and hence lacks global planning for the future. To address this problem, we introduce another decoder, called seer decoder, into the encoder-decoder framework during training, which involves future information in target predictions. Meanwhile, we force the conventional decoder to simulate the behaviors of the seer decoder via knowledge distillation. In this way, at test the conventional decoder can perform like the seer decoder without the attendance of it. Experiment results on the Chinese-English, English-German and English-Romanian translation tasks show our method can outperform competitive baselines significantly and achieves greater improvements on the bigger data sets. Besides, the experiments also prove knowledge distillation the best way to transfer knowledge from the seer decoder to the conventional decoder compared to adversarial learning and L2 regularization.

📄 PDF Abstract BibTeX arXiv:2106.06751

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderKnowledge DistillationL2 RegularizationMachine TranslationTranslation

Methods 이 논문이 사용한 방법론

Knowledge Distillation A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions.…

Similar Papers 제목 키워드 기반

Parallel Attention Forcing for Machine Translation

2022-11-06 · Qingyun Dou, Mark Gales

Attention-based autoregressive models have achieved state-of-the-art performance in various sequence-to-sequence tasks, including Text-To-Speech (TTS) and Neural Machine Translation (NMT), but can be difficult to train. …

Machine TranslationNMTtext-to-speechText to Speech+1

Attention Forcing for Sequence-to-sequence Model Training

2019-09-26 · Qingyun Dou, Yiting Lu, Joshua Efiong, Mark J. F. Gales

Auto-regressive sequence-to-sequence models with attention mechanism have achieved state-of-the-art performance in many tasks such as machine translation and speech synthesis. These models can be difficult to train. The …

Machine TranslationmodelSpeech SynthesisTranslation+1

TeaForN: Teacher-Forcing with N-grams

2020-10-07 · EMNLP 2020 11 · Sebastian Goodman, Nan Ding, Radu Soricut

Sequence generation models trained with teacher-forcing suffer from issues related to exposure bias and lack of differentiability across timesteps. Our proposed method, Teacher-Forcing with N-grams (TeaForN), addresses b…

DecoderMachine TranslationNews SummarizationTranslation

Context Forcing: Consistent Autoregressive Video Generation with Long Context

2026-02-05 · Shuo Chen, Cong Wei, Sun Sun, Ping Nie 외 arxiv

Recent approaches to real-time long video generation typically employ streaming tuning strategies, attempting to train a long-context student using a short-context (memoryless) teacher. In these frameworks, the student p…

Video Generation

Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation

2026-02-02 · Hongzhou Zhu, Min Zhao, Guande He, Hang Su 외 arxiv

To achieve real-time interactive video generation, current methods distill pretrained bidirectional video diffusion models into few-step autoregressive (AR) models, facing an architectural gap when full attention is repl…

Instruction FollowingVideo Generation