paper-with-me

홈 › Papers

Independent language modeling architecture for end-to-end ASR

2019-11-25 · Van Tung Pham, Hai-Hua Xu, Yerbolat Khassanov, Zhiping Zeng, Eng Siong Chng, Chongjia Ni, Bin Ma, Haizhou Li

The attention-based end-to-end (E2E) automatic speech recognition (ASR) architecture allows for joint optimization of acoustic and language models within a single network. However, in a vanilla E2E ASR architecture, the decoder sub-network (subnet), which incorporates the role of the language model (LM), is conditioned on the encoder output. This means that the acoustic encoder and the language model are entangled that doesn't allow language model to be trained separately from external text data. To address this problem, in this work, we propose a new architecture that separates the decoder subnet from the encoder output. In this way, the decoupled subnet becomes an independently trainable LM subnet, which can easily be updated using the external text data. We study two strategies for updating the new architecture. Experimental results show that, 1) the independent LM architecture benefits from external text data, achieving 9.3% and 22.8% relative character and word error rate reduction on Mandarin HKUST and English NSC datasets respectively; 2)the proposed architecture works well with external LM and can be generalized to different amount of labelled data.

📄 PDF Abstract BibTeX arXiv:1912.00863

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderLanguage ModelingLanguage Modellingspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

The Dual-Stream Transformer: Channelized Architecture for Interpretable Language Modeling

2026-03-08 · J. Clayton Kerce, Alexis Fox arxiv

Standard transformers entangle all computation in a single residual stream, obscuring which components perform which functions. We introduce the Dual-Stream Transformer, which decomposes the residual stream into two func…

GQ-VAE: A gated quantized VAE for learning variable length tokens

2025-12-26 · Theo Datta, Kayla Huang, Sham Kakade, David Brandfonbrener arxiv

While most frontier models still use deterministic frequency-based tokenization algorithms such as byte-pair encoding (BPE), there has been significant recent work to design learned neural tokenizers. However, these sche…

On the Relation between Linguistic Typology and (Limitations of) Multilingual Language Modeling

2018-10-01 · EMNLP 2018 10 · Daniela Gerz, Ivan Vuli{\'c}, Edoardo Maria Ponti, Roi Reichart 외

A key challenge in cross-lingual NLP is developing general language-independent architectures that are equally applicable to any language. However, this ambition is largely hampered by the variation in structural and sem…

Language ModelingLanguage ModellingRelation

Conversation Modeling on Reddit using a Graph-Structured LSTM

2017-04-07 · TACL 2018 1 · Vicky Zayats, Mari Ostendorf

This paper presents a novel approach for modeling threaded discussions on social media using a graph-structured bidirectional LSTM which represents both hierarchical and temporal conversation structure. In experiments wi…

ÚFAL at MRP 2020: Permutation-invariant Semantic Parsing in PERIN

2020-11-02 · David Samuel, Milan Straka

We present PERIN, a novel permutation-invariant approach to sentence-to-graph semantic parsing. PERIN is a versatile, cross-framework and language independent architecture for universal modeling of semantic structures. O…

Semantic ParsingSentence