paper-with-me

홈 › Papers

Attention-based Vocabulary Selection for NMT Decoding

2017-06-12 · Baskaran Sankaran, Markus Freitag, Yaser Al-Onaizan

Neural Machine Translation (NMT) models usually use large target vocabulary sizes to capture most of the words in the target language. The vocabulary size is a big factor when decoding new sentences as the final softmax layer normalizes over all possible target words. To address this problem, it is widely common to restrict the target vocabulary with candidate lists based on the source sentence. Usually, the candidate lists are a combination of external word-to-word aligner, phrase table entries or most frequent words. In this work, we propose a simple and yet novel approach to learn candidate lists directly from the attention layer during NMT training. The candidate lists are highly optimized for the current NMT model and do not need any external computation of the candidate pool. We show significant decoding speedup compared with using the entire vocabulary, without losing any translation quality for two language pairs.

📄 PDF Abstract BibTeX arXiv:1706.03824

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationNMTSentenceTranslation

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…

Similar Papers 제목 키워드 기반

Attention-Guided Layer Selection for Contrastive Decoding in Large Language Models

2026-07-25 · Yusuke Sakai, Natthawut Kertkeidkachorn, Kiyoaki Shirai arxiv

Contrastive decoding methods such as DoLa improve the factuality of Large Language Models (LLMs) by contrasting the output distributions of mature and premature layers. However, DoLa's dynamic layer selection relies sole…

Balancing Coverage and Draft Latency in Vocabulary Trimming for Faster Speculative Decoding

2026-03-05 · Ofir Ben Shoham arxiv

Speculative decoding accelerates inference for Large Language Models by using a lightweight draft model to propose candidate tokens that are verified in parallel by a larger target model. Prior work shows that the draft …

Inferring from Logits: Exploring Best Practices for Decoding-Free Generative Candidate Selection

2025-01-28 · Mingyu Derek Ma, Yanna Ding, Zijie Huang, Jianxi Gao 외

Generative Language Models rely on autoregressive decoding to produce the output sequence token by token. Many tasks such as preference optimization, require the model to produce task-level output consisting of multiple …

Multiple-choice

Enabling Real-time Neural IME with Incremental Vocabulary Selection

2019-06-01 · NAACL 2019 6 · Jiali Yao, Raphael Shu, Xinjian Li, Katsutoshi Ohtsuki 외

Input method editor (IME) converts sequential alphabet key inputs to words in a target language. It is an indispensable service for billions of Asian users. Although the neural-based language model is extensively studied…

CPULanguage ModelingLanguage Modellingspeech-recognition+1

EvoSpec: Evolving Speculative Decoding via Real-Time Vocabulary and Parameter Adaptation

2026-04-17 · Shuyu Zhang, Lingfeng Pan, Qicheng Wang, Yaqi Shi 외 arxiv

Speculative decoding accelerates Large Language Model inference through draft-then-verify generation, yet lightweight draft models face coupled efficiency and quality limitations: large-vocabulary output projection is co…