paper-with-me

홈 › Papers

Input Combination Strategies for Multi-Source Transformer Decoder

2018-11-12 · Jindřich Libovický, Jindřich Helcl, David Mareček

In multi-source sequence-to-sequence tasks, the attention mechanism can be modeled in several ways. This topic has been thoroughly studied on recurrent architectures. In this paper, we extend the previous work to the encoder-decoder attention in the Transformer architecture. We propose four different input combination strategies for the encoder-decoder attention: serial, parallel, flat, and hierarchical. We evaluate our methods on tasks of multimodal translation and translation with multiple source languages. The experiments show that the models are able to use multiple sources and improve over single source baselines.

📄 PDF Abstract BibTeX arXiv:1811.04716

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderTranslation

Methods 이 논문이 사용한 방법론

Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Position-Wise Feed-Forward Layer 설명 없음
Residual Connection 설명 없음
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…

Similar Papers 제목 키워드 기반

Input Combination Strategies for Multi-Source Transformer Decoder

2018-10-01 · WS 2018 10 · Jind{\v{r}}ich Libovick{\'y}, Jind{\v{r}}ich Helcl, David Mare{\v{c}}ek

In multi-source sequence-to-sequence tasks, the attention mechanism can be modeled in several ways. This topic has been thoroughly studied on recurrent architectures. In this paper, we extend the previous work to the enc…

DecoderImage CaptioningMachine TranslationMultimodal Machine Translation+2

MVP: Multi-source Voice Pathology detection

2025-05-26 · Alkis Koudounas, Moreno La Quatra, Gabriele Ciravegna, Marco Fantini 외

Voice disorders significantly impact patient quality of life, yet non-invasive automated diagnosis remains under-explored due to both the scarcity of pathological voice data, and the variability in recording sources. Thi…

SentenceVoice pathology detection

Fine-tuning Transformers with Additional Context to Classify Discursive Moves in Mathematics Classrooms

2022-07-01 · NAACL (BEA) 2022 7 · Abhijit Suresh, Jennifer Jacobs, Margaret Perkoff, James H. Martin 외

“Talk moves” are specific discursive strategies used by teachers and students to facilitate conversations in which students share their thinking, and actively consider the ideas of others, and engage in rich discussions.…

Data Augmentation for Transformer-based G2P

2020-07-01 · WS 2020 7 · Zach Ryan, Mans Hulden

The Transformer model has been shown to outperform other neural seq2seq models in several character-level tasks. It is unclear, however, if the Transformer would benefit as much as other seq2seq models from data augmenta…

Data Augmentation

Automatic Discovery of Composite SPMD Partitioning Strategies in PartIR

2022-10-07 · Sami Alabed, Dominik Grewe, Juliana Franco, Bart Chrzaszcz 외

Large neural network models are commonly trained through a combination of advanced parallelism strategies in a single program, multiple data (SPMD) paradigm. For example, training large transformer models requires combin…