paper-with-me

Papers

Efficient Long-Form Speech Recognition for General Speech In-Context Learning

2024-09-29 · Hao Yen, Shaoshi Ling, Guoli Ye

We propose a novel approach to end-to-end automatic speech recognition (ASR) to achieve efficient speech in-context learning (SICL) for (i) long-form speech decoding, (ii) test-time speaker adaptation, and (iii) test-time contextual biasing. Specifically, we introduce an attention-based encoder-decoder (AED) model with SICL capability (referred to as SICL-AED), where the decoder utilizes an utterance-level cross-attention to integrate information from the encoder's output efficiently, and a document-level self-attention to learn contextual information. Evaluated on the benchmark TEDLIUM3 dataset, SICL-AED achieves an 8.64% relative word error rate (WER) reduction compared to a baseline utterance-level AED model by leveraging previously decoded outputs as in-context examples. It also demonstrates comparable performance to conventional long-form AED systems with significantly reduced runtime and memory complexity. Additionally, we introduce an in-context fine-tuning (ICFT) technique that further enhances SICL effectiveness during inference. Experiments on speaker adaptation and contextual biasing highlight the general speech in-context learning capabilities of our system, achieving effective results with provided contexts. Without specific fine-tuning, SICL-AED matches the performance of supervised AED baselines for speaker adaptation and improves entity recall by 64% for contextual biasing task.

📄 PDF Abstract BibTeX arXiv:2409.19757

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderFormIn-Context Learningspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Inner speech recognition through electroencephalographic signals

2022-10-11 · Francesca Gasparini, Elisa Cazzaniga, Aurora Saibene

This work focuses on inner speech recognition starting from EEG signals. Inner speech recognition is defined as the internalized process in which the person thinks in pure meanings, generally associated with an auditory …

EEGElectroencephalogram (EEG)speech-recognitionSpeech Recognition

SememeASR: Boosting Performance of End-to-End Speech Recognition against Domain and Long-Tailed Data Shift with Sememe Semantic Knowledge

2023-09-04 · Jiaxu Zhu, Changhe Song, Zhiyong Wu, Helen Meng

Recently, excellent progress has been made in speech recognition. However, pure data-driven approaches have struggled to solve the problem in domain-mismatch and long-tailed data. Considering that knowledge-driven approa…

Domain Generalizationspeech-recognitionSpeech Recognition

Lombard Effect for Bilingual Speakers in Cantonese and English: importance of spectro-temporal features

2022-04-14 · Maximilian Karl Scharf, Sabine Hochmuth, Lena L. N. Wong, Birger Kollmeier 외

For a better understanding of the mechanisms underlying speech perception and the contribution of different signal features, computational models of speech recognition have a long tradition in hearing research. Due to th…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Analysis and Tuning of a Voice Assistant System for Dysfluent Speech

2021-06-18 · Vikramjit Mitra, Zifang Huang, Colin Lea, Lauren Tooley 외

Dysfluencies and variations in speech pronunciation can severely degrade speech recognition performance, and for many individuals with moderate-to-severe speech disorders, voice operated systems do not work. Current spee…

Intent Recognitionspeech-recognitionSpeech Recognition

Directed Speech Separation for Automatic Speech Recognition of Long Form Conversational Speech

2021-12-10 · Rohit Paturi, Sundararajan Srinivasan, Katrin Kirchhoff, Daniel Garcia-Romero

Many of the recent advances in speech separation are primarily aimed at synthetic mixtures of short audio utterances with high degrees of overlap. Most of these approaches need an additional stitching step to stitch the …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Formspeech-recognition+2