paper-with-me

홈 › Papers

MTLM: Incorporating Bidirectional Text Information to Enhance Language Model Training in Speech Recognition Systems

2025-02-14 · Qingliang Meng, Pengju Ren, Tian Li, Changsong Dai, HuiZhi Liang

Automatic speech recognition (ASR) systems normally consist of an acoustic model (AM) and a language model (LM). The acoustic model estimates the probability distribution of text given the input speech, while the language model calibrates this distribution toward a specific knowledge domain to produce the final transcription. Traditional ASR-specific LMs are typically trained in a unidirectional (left-to-right) manner to align with autoregressive decoding. However, this restricts the model from leveraging the right-side context during training, limiting its representational capacity. In this work, we propose MTLM, a novel training paradigm that unifies unidirectional and bidirectional manners through 3 training objectives: ULM, BMLM, and UMLM. This approach enhances the LM's ability to capture richer linguistic patterns from both left and right contexts while preserving compatibility with standard ASR autoregressive decoding methods. As a result, the MTLM model not only enhances the ASR system's performance but also support multiple decoding strategies, including shallow fusion, unidirectional/bidirectional n-best rescoring. Experiments on the LibriSpeech dataset show that MTLM consistently outperforms unidirectional training across multiple decoding strategies, highlighting its effectiveness and flexibility in ASR applications.

📄 PDF Abstract BibTeX arXiv:2502.10058

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modellingspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Bidirectional Trained Tree-Structured Decoder for Handwritten Mathematical Expression Recognition

2023-12-31 · Hanbo Cheng, Chenyu Liu, Pengfei Hu, Zhenrong Zhang 외

The Handwritten Mathematical Expression Recognition (HMER) task is a critical branch in the field of OCR. Recent studies have demonstrated that incorporating bidirectional context information significantly improves the p…

DecoderLanguage ModelingLanguage ModellingOptical Character Recognition (OCR)

TemPose-TF-ASF: Two-Stage Bidirectional Stroke Context Fusion for Badminton Stroke Classification

2026-05-04 · Tzu-Yu Liu, Duan-Shin Lee arxiv

Accurate badminton stroke prediction is crucial for fine-grained sports analysis and tactical decision support. However, existing methods struggle to model rich temporal context. This paper introduces TemPose-TF-ASF (Adj…

Stroke Classification

TiBiX: Leveraging Temporal Information for Bidirectional X-ray and Report Generation

2024-03-20 · Santosh Sanjeev, Fadillah Adamsyah Maani, Arsen Abzhanov, Vijay Ram Papineni 외

With the emergence of vision language models in the medical imaging domain, numerous studies have focused on two dominant research activities: (1) report generation from Chest X-rays (CXR), and (2) synthetic scan generat…

Image Generation

Cross-Modal Bidirectional Interaction Model for Referring Remote Sensing Image Segmentation

2024-10-11 · Zhe Dong, Yuzhe Sun, Yanfeng Gu, Tianzhu Liu

Given a natural language expression and a remote sensing image, the goal of referring remote sensing image segmentation (RRSIS) is to generate a pixel-level mask of the target object identified by the referring expressio…

BenchmarkingImage SegmentationReferring ExpressionSemantic Segmentation

MSD-KMamba: Bidirectional Spatial-Aware Multi-Modal 3D Brain Segmentation via Multi-scale Self-Distilled Fusion Strategy

2025-09-28 · Dayu Tan, Ziwei Zhang, Yansan Su, Xin Peng 외 arxiv

Numerous CNN-Transformer hybrid models rely on high-complexity global attention mechanisms to capture long-range dependencies, which introduces non-linear computational complexity and leads to significant resource consum…

Computational EfficiencyKnowledge DistillationBrain SegmentationImage Segmentation