paper-with-me

Papers

Swedish Whispers; Leveraging a Massive Speech Corpus for Swedish Speech Recognition

2025-05-23 · Leonora Vesterbacka, Faton Rekathati, Robin Kurtz, Justyna Sikora, Agnes Toftgård

This work presents a suite of fine-tuned Whisper models for Swedish, trained on a dataset of unprecedented size and variability for this mid-resourced language. As languages of smaller sizes are often underrepresented in multilingual training datasets, substantial improvements in performance can be achieved by fine-tuning existing multilingual models, as shown in this work. This work reports an overall improvement across model sizes compared to OpenAI's Whisper evaluated on Swedish. Most notably, we report an average 47% reduction in WER comparing our best performing model to OpenAI's whisper-large-v3, in evaluations across FLEURS, Common Voice, and NST.

📄 PDF Abstract BibTeX arXiv:2505.17538

Code (0)

등록된 구현이 없습니다.

Tasks

speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

SL\"aNDa: An Annotated Corpus of Narrative and Dialogue in Swedish Literary Fiction

2020-05-01 · LREC 2020 5 · Sara Stymne, Carin {\"O}stman

We describe a new corpus, SL{\"a}NDa, the Swedish Literary corpus of Narrative and Dialogue. It contains Swedish literary fiction, which has been manually annotated for cited materials, with a focus on dialogue. The anno…

Hearing voices at the National Library -- a speech corpus and acoustic model for the Swedish language

2022-05-06 · Martin Malmsten, Chris Haffenden, Love Börjeson

This paper explains our work in developing new acoustic models for automated speech recognition (ASR) at KBLab, the infrastructure for data-driven research at the National Library of Sweden (KB). We evaluate different ap…

speech-recognitionSpeech RecognitionSpeech-to-Text

Chinese Whispers: A Multimodal Dataset for Embodied Language Grounding

2020-05-01 · LREC 2020 5 · Dimosthenis Kontogiorgos, Elena Sibirtseva, Joakim Gustafson

In this paper, we introduce a multimodal dataset in which subjects are instructing each other how to assemble IKEA furniture. Using the concept of {`}Chinese Whispers{'}, an old children{'}s game, we employ a novel metho…

Advancing NAM-to-Speech Conversion with Novel Methods and the MultiNAM Dataset

2024-12-25 · Neil Shah, Shirish Karande, Vineet Gandhi

Current Non-Audible Murmur (NAM)-to-speech techniques rely on voice cloning to simulate ground-truth speech from paired whispers. However, the simulated speech often lacks intelligibility and fails to generalize well acr…

text-to-speechText to SpeechVoice Cloning

Enriching the Swedish Sign Language Corpus with Part of Speech Tags Using Joint Bayesian Word Alignment and Annotation Transfer

2015-05-01 · WS 2015 5 · Robert {\"O}stling, Carl B{\"o}rstell, Lars Wallin
Word Alignment