paper-with-me

Papers

Token-Level Serialized Output Training for Joint Streaming ASR and ST Leveraging Textual Alignments

2023-07-07 · Sara Papi, Peidong Wang, Junkun Chen, Jian Xue, Jinyu Li, Yashesh Gaur

In real-world applications, users often require both translations and transcriptions of speech to enhance their comprehension, particularly in streaming scenarios where incremental generation is necessary. This paper introduces a streaming Transformer-Transducer that jointly generates automatic speech recognition (ASR) and speech translation (ST) outputs using a single decoder. To produce ASR and ST content effectively with minimal latency, we propose a joint token-level serialized output training method that interleaves source and target words by leveraging an off-the-shelf textual aligner. Experiments in monolingual (it-en) and multilingual (\{de,es,it\}-en) settings demonstrate that our approach achieves the best quality-latency balance. With an average ASR latency of 1s and ST latency of 1.3s, our model shows no degradation or even improves output quality compared to separate ASR and ST models, yielding an average improvement of 1.1 WER and 0.4 BLEU in the multilingual case.

📄 PDF Abstract BibTeX arXiv:2307.03354

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Decoderspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Leveraging Timestamp Information for Serialized Joint Streaming Recognition and Translation

2023-10-23 · Sara Papi, Peidong Wang, Junkun Chen, Jian Xue 외

The growing need for instant spoken language transcription and translation is driven by increased global communication and cross-lingual interactions. This has made offering translations in multiple languages essential f…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Decoderde-en+3

BA-SOT: Boundary-Aware Serialized Output Training for Multi-Talker ASR

2023-05-23 · Yuhao Liang, Fan Yu, Yangze Li, Pengcheng Guo 외

The recently proposed serialized output training (SOT) simplifies multi-talker automatic speech recognition (ASR) by generating speaker transcriptions separated by a special token. However, frequent speaker changes can m…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Change DetectionDecoder+2

Streaming Multi-Talker ASR with Token-Level Serialized Output Training

2022-02-02 · Naoyuki Kanda, Jian Wu, Yu Wu, Xiong Xiao 외

This paper proposes a token-level serialized output training (t-SOT), a novel framework for streaming multi-talker automatic speech recognition (ASR). Unlike existing streaming multi-talker ASR models using multiple outp…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Joint ASR and Speaker Role Tagging with Serialized Output Training

2025-06-12 · Anfeng Xu, Tiantian Feng, Shrikanth Narayanan

Automatic Speech Recognition systems have made significant progress with large-scale pre-trained models. However, most current systems focus solely on transcribing the speech without identifying speaker roles, a function…

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

Streaming Speaker-Attributed ASR with Token-Level Speaker Embeddings

2022-03-30 · Naoyuki Kanda, Jian Wu, Yu Wu, Xiong Xiao 외

This paper presents a streaming speaker-attributed automatic speech recognition (SA-ASR) model that can recognize ``who spoke what'' with low latency even when multiple people are speaking simultaneously. Our model is ba…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Decoderspeaker-diarization+4