paper-with-me

Papers

Leveraging Timestamp Information for Serialized Joint Streaming Recognition and Translation

2023-10-23 · Sara Papi, Peidong Wang, Junkun Chen, Jian Xue, Naoyuki Kanda, Jinyu Li, Yashesh Gaur

The growing need for instant spoken language transcription and translation is driven by increased global communication and cross-lingual interactions. This has made offering translations in multiple languages essential for user applications. Traditional approaches to automatic speech recognition (ASR) and speech translation (ST) have often relied on separate systems, leading to inefficiencies in computational resources, and increased synchronization complexity in real time. In this paper, we propose a streaming Transformer-Transducer (T-T) model able to jointly produce many-to-one and one-to-many transcription and translation using a single decoder. We introduce a novel method for joint token-level serialized output training based on timestamp information to effectively produce ASR and ST outputs in the streaming setting. Experiments on {it,es,de}->en prove the effectiveness of our approach, enabling the generation of one-to-many joint outputs with a single decoder for the first time.

📄 PDF Abstract BibTeX arXiv:2310.14806

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Decoderde-enspeech-recognitionSpeech RecognitionTranslation

Similar Papers 제목 키워드 기반

Token-Level Serialized Output Training for Joint Streaming ASR and ST Leveraging Textual Alignments

2023-07-07 · Sara Papi, Peidong Wang, Junkun Chen, Jian Xue 외

In real-world applications, users often require both translations and transcriptions of speech to enhance their comprehension, particularly in streaming scenarios where incremental generation is necessary. This paper int…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Decoderspeech-recognition+1

Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition

2025-10-04 · Martin Kocour, Martin Karafiat, Alexander Polok, Dominik Klement 외 arxiv

We propose a speaker-attributed (SA) Whisper-based model for multi-talker speech recognition that combines target-speaker modeling with serialized output training (SOT). Our approach leverages a Diarization-Conditioned W…

Speech Recognition

Adapting Multi-Lingual ASR Models for Handling Multiple Talkers

2023-05-30 · Chenda Li, Yao Qian, Zhuo Chen, Naoyuki Kanda 외

State-of-the-art large-scale universal speech models (USMs) show a decent automatic speech recognition (ASR) performance across multiple domains and languages. However, it remains a challenge for these models to recogniz…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Streaming Speaker-Attributed ASR with Token-Level Speaker Embeddings

2022-03-30 · Naoyuki Kanda, Jian Wu, Yu Wu, Xiong Xiao 외

This paper presents a streaming speaker-attributed automatic speech recognition (SA-ASR) model that can recognize ``who spoke what'' with low latency even when multiple people are speaking simultaneously. Our model is ba…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Decoderspeaker-diarization+4

Speaker-Distinguishable CTC: Learning Speaker Distinction Using CTC for Multi-Talker Speech Recognition

2025-06-09 · Asahi Sakuma, Hiroaki Sato, Ryuga Sugano, Tadashi Kumano 외

This paper presents a novel framework for multi-talker automatic speech recognition without the need for auxiliary information. Serialized Output Training (SOT), a widely used approach, suffers from recognition errors du…

Automatic Speech RecognitionMulti-Task Learningspeech-recognitionSpeech Recognition