paper-with-me

홈 › Papers

SPGISpeech 2.0: Transcribed multi-speaker financial audio for speaker-tagged transcription

2025-08-07 · Raymond Grossman, Taejin Park, Kunal Dhawan, Andrew Titus, Sophia Zhi, Yulia Shchadilova, Weiqing Wang, Jagadeesh Balam, Boris Ginsburg arxiv

We introduce SPGISpeech 2.0, a dataset suitable for speaker-tagged transcription in the financial domain. SPGISpeech 2.0 improves the diversity of applicable modeling tasks while maintaining the core characteristic of the original SPGISpeech dataset: audio snippets and their corresponding fully formatted text transcriptions, usable for end-to-end automatic speech recognition (ASR). SPGISpeech 2.0 consists of 3,780 additional hours of professionally transcribed earnings calls. Furthermore, the dataset contains call and speaker information for each audio snippet facilitating multi-talker ASR. We validate the utility of SPGISpeech 2.0 through improvements in speaker-tagged ASR performance of popular speech recognition models after fine-tuning on SPGISpeech 2.0. Released free for non-commercial use, we expect SPGISpeech 2.0 to foster advancements in speech recognition technologies and inspire a wide range of research applications.

📄 PDF Abstract BibTeX arXiv:2508.05554

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Recognition

Similar Papers 제목 키워드 기반

SPGISpeech: 5,000 hours of transcribed financial audio for fully formatted end-to-end speech recognition

2021-04-05 · Patrick K. O'Neill, Vitaly Lavrukhin, Somshubra Majumdar, Vahid Noroozi 외

In the English speech-to-text (STT) machine learning task, acoustic models are conventionally trained on uncased Latin characters, and any necessary orthography (such as capitalization, punctuation, and denormalization o…

speech-recognitionSpeech RecognitionSpeech-to-Text

Multimodal speech synthesis architecture for unsupervised speaker adaptation

2018-08-20 · Hieu-Thi Luong, Junichi Yamagishi

This paper proposes a new architecture for speaker adaptation of multi-speaker neural-network speech synthesis systems, in which an unseen speaker's voice can be built using a relatively small amount of speech data witho…

Speech Synthesis

Fitting New Speakers Based on a Short Untranscribed Sample

2018-02-20 · ICML 2018 7 · Eliya Nachmani, Adam Polyak, Yaniv Taigman, Lior Wolf

Learning-based Text To Speech systems have the potential to generalize from one speaker to the next and thus require a relatively short sample of any new voice. However, this promise is currently largely unrealized. We p…

Speech Synthesistext-to-speechText to Speech

Adversarial Speaker-Consistency Learning Using Untranscribed Speech Data for Zero-Shot Multi-Speaker Text-to-Speech

2022-10-12 · Byoung Jin Choi, Myeonghun Jeong, Minchan Kim, Sung Hwan Mun 외

Several recently proposed text-to-speech (TTS) models achieved to generate the speech samples with the human-level quality in the single-speaker and multi-speaker TTS scenarios with a set of pre-defined speakers. However…

text-to-speechText to Speech

Exact Prosody Cloning in Zero-Shot Multispeaker Text-to-Speech

2022-06-24 · Florian Lux, Julia Koch, Ngoc Thang Vu

The cloning of a speaker's voice using an untranscribed reference sample is one of the great advances of modern neural text-to-speech (TTS) methods. Approaches for mimicking the prosody of a transcribed reference audio h…

text-to-speechText to Speech