paper-with-me

홈 › Papers

Sequence Transduction with Graph-based Supervision

2021-11-01 · Niko Moritz, Takaaki Hori, Shinji Watanabe, Jonathan Le Roux

The recurrent neural network transducer (RNN-T) objective plays a major role in building today's best automatic speech recognition (ASR) systems for production. Similarly to the connectionist temporal classification (CTC) objective, the RNN-T loss uses specific rules that define how a set of alignments is generated to form a lattice for the full-sum training. However, it is yet largely unknown if these rules are optimal and do lead to the best possible ASR results. In this work, we present a new transducer objective function that generalizes the RNN-T loss to accept a graph representation of the labels, thus providing a flexible and efficient framework to manipulate training lattices, e.g., for studying different transition rules, implementing different transducer losses, or restricting alignments. We demonstrate that transducer-based ASR with CTC-like lattice achieves better results compared to standard RNN-T, while also ensuring a strictly monotonic alignment, which will allow better optimization of the decoding procedure. For example, the proposed CTC-like transducer achieves an improvement of 4.8% on the test-other condition of LibriSpeech relative to an equivalent RNN-T based system.

📄 PDF Abstract BibTeX arXiv:2111.01272

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

AMR Parsing as Sequence-to-Graph Transduction

2019-05-21 · ACL 2019 7 · Sheng Zhang, Xutai Ma, Kevin Duh, Benjamin Van Durme

We propose an attention-based model that treats AMR parsing as sequence-to-graph transduction. Unlike most AMR parsers that rely on pre-trained aligners, external semantic resources, or data augmentation, our proposed pa…

AMR ParsingSemantic Parsing

Exact Hard Monotonic Attention for Character-Level Transduction

2019-05-15 · ACL 2019 7 · Shijie Wu, Ryan Cotterell

Many common character-level, string-to string transduction tasks, e.g., grapheme-tophoneme conversion and morphological inflection, consist almost exclusively of monotonic transductions. However, neural sequence-to seque…

Hard AttentionInductive BiasMorphological Inflection

String Transduction with Target Language Models and Insertion Handling

2018-09-19 · WS 2018 10 · Garrett Nicolai, Saeed Najafi, Grzegorz Kondrak

Many character-level tasks can be framed as sequence-to-sequence transduction, where the target is a word from a natural language. We show that leveraging target language models derived from unannotated target corpora, c…

Applying the Transformer to Character-level Transduction

2020-05-20 · EACL 2021 2 · Shijie Wu, Ryan Cotterell, Mans Hulden

The transformer has been shown to outperform recurrent neural network-based sequence-to-sequence models in various word-level NLP tasks. Yet for character-level transduction tasks, e.g. morphological inflection generatio…

Grapheme-to-Phoneme ConversionMorphological InflectionText NormalizationTransliteration

Sequence Transduction with Recurrent Neural Networks

2012-11-14 · Alex Graves

Many machine learning tasks can be expressed as the transformation---or \emph{transduction}---of input sequences into output sequences: speech recognition, machine translation, protein secondary structure prediction and …

Machine TranslationPhoneme RecognitionSpeech Recognitiontext-to-speech+2