paper-with-me

홈 › Papers

Internal Language Model Estimation Through Explicit Context Vector Learning for Attention-based Encoder-decoder ASR

2022-01-26 · Yufei Liu, Rao Ma, HaiHua Xu, Yi He, Zejun Ma, Weibin Zhang

An end-to-end (E2E) ASR model implicitly learns a prior Internal Language Model (ILM) from the training transcripts. To fuse an external LM using Bayes posterior theory, the log likelihood produced by the ILM has to be accurately estimated and subtracted. In this paper we propose two novel approaches to estimate the ILM based on Listen-Attend-Spell (LAS) framework. The first method is to replace the context vector of the LAS decoder at every time step with a vector that is learned with training transcripts. Furthermore, we propose another method that uses a lightweight feed-forward network to directly map query vector to context vector in a dynamic sense. Since the context vectors are learned by minimizing the perplexities on training transcripts, and their estimation is independent of encoder output, hence the ILMs are accurately learned for both methods. Experiments show that the ILMs achieve the lowest perplexity, indicating the efficacy of the proposed methods. In addition, they also significantly outperform the shallow fusion method, as well as two previously proposed ILM Estimation (ILME) approaches on several datasets.

📄 PDF Abstract BibTeX arXiv:2201.11627

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderLanguage ModelingLanguage ModellingSpeech Recognition

Similar Papers 제목 키워드 기반

Beyond Static Perception: Integrating Temporal Context into VLMs for Cloth Folding

2025-05-12 · Oriol Barbany, Adrià Colomé, Carme Torras

Manipulating clothes is challenging due to their complex dynamics, high deformability, and frequent self-occlusions. Garments exhibit a nearly infinite number of configurations, making explicit state representations diff…

State Estimation

Navigating Unreliable Parametric and Contextual Knowledge: Explicit Knowledge Conflict Resolution for LLM Inference

2026-06-18 · Huang Peng, Jiuyang Tang, Weixin Zeng, Hao Xu 외 arxiv

Large language models (LLMs) have achieved strong performance across a wide range of language-based tasks by leveraging both extensive parametric knowledge and in-context learning ability, enabling them to incorporate ex…

Label-Context-Dependent Internal Language Model Estimation for CTC

2025-06-06 · Zijian Yang, Minh-Nghia Phan, Ralf Schlüter, Hermann Ney

Although connectionist temporal classification (CTC) has the label context independence assumption, it can still implicitly learn a context-dependent internal language model (ILM) due to modern powerful encoders. In this…

Knowledge DistillationLanguage ModelingLanguage Modelling

InternalInspector $I^2$: Robust Confidence Estimation in LLMs through Internal States

2024-06-17 · Mohammad Beigi, Ying Shen, Runing Yang, Zihao Lin 외

Despite their vast capabilities, Large Language Models (LLMs) often struggle with generating reliable outputs, frequently producing high-confidence inaccuracies known as hallucinations. Addressing this challenge, our res…

BenchmarkingContrastive LearningHallucinationNatural Language Understanding+2

Beyond Chains of Thought: Benchmarking Latent-Space Reasoning Abilities in Large Language Models

2025-04-14 · Thilo Hagendorff, Sarah Fabi

Large language models (LLMs) can perform reasoning computations both internally within their latent space and externally by generating explicit token sequences like chains of thought. Significant progress in enhancing re…

BenchmarkingDescriptive