paper-with-me

홈 › Papers

TOGGL: Transcribing Overlapping Speech with Staggered Labeling

2024-08-12 · Chak-Fai Li, William Hartmann, Matthew Snover

Transcribing the speech of multiple overlapping speakers typically requires separating the audio into multiple streams and recognizing each one independently. More recent work jointly separates and transcribes, but requires a separate decoding component for each speaker. We propose the TOGGL model to simultaneously transcribe the speech of multiple speakers. The TOGGL model uses special output tokens to attribute the speech to each speaker with only a single decoder. Our approach generalizes beyond two speakers, even when trained only on two-speaker data. We demonstrate superior performance compared to competing approaches on a conversational speech dataset. Our approach also improves performance on single-speaker audio.

📄 PDF Abstract BibTeX arXiv:2408.06474

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeDecoder

Similar Papers 제목 키워드 기반

Speaker-Reasoner: Scaling Interaction Turns and Reasoning Patterns for Timestamped Speaker-Attributed ASR

2026-04-03 · Zhennan Lin, Shuai Wang, Zhaokai Sun, Pengyuan Xie 외 arxiv

Transcribing and understanding multi-speaker conversations requires speech recognition, speaker attribution, and timestamp localization. While speech LLMs excel at single-speaker tasks, multi-speaker scenarios remain cha…

Speech Recognition

Introducing a web application for labeling, visualizing speech and correcting derived speech signals

2014-05-01 · LREC 2014 5 · Raphael Winkelmann, Georg Raess

The advent of HTML5 has sparked a great increase in interest in the web as a development platform for a variety of different research applications. Due to its ability to easily deploy software to remote clients and the r…

Management

Low-Latency Speaker-Independent Continuous Speech Separation

2019-04-13 · Takuya Yoshioka, Zhuo Chen, Changliang Liu, Xiong Xiao 외

Speaker independent continuous speech separation (SI-CSS) is a task of converting a continuous audio stream, which may contain overlapping voices of unknown speakers, into a fixed number of continuous signals each of whi…

speech-recognitionSpeech RecognitionSpeech Separation

Large scale weakly and semi-supervised learning for low-resource video ASR

2020-05-16 · Kritika Singh, Vimal Manohar, Alex Xiao, Sergey Edunov 외

Many semi- and weakly-supervised approaches have been investigated for overcoming the labeling cost of building high quality speech recognition systems. On the challenging task of transcribing social media videos in low-…

Decoderspeech-recognitionSpeech Recognition

Contrastive Semi-supervised Learning for ASR

2021-03-09 · Alex Xiao, Christian Fuegen, Abdelrahman Mohamed

Pseudo-labeling is the most adopted method for pre-training automatic speech recognition (ASR) models. However, its performance suffers from the supervised teacher model's degrading quality in low-resource setups and und…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Representation Learningspeech-recognition+1