paper-with-me

Papers

Adapting OpenAI's Whisper for Speech Recognition on Code-Switch Mandarin-English SEAME and ASRU2019 Datasets

2023-11-29 · Yuhang Yang, Yizhou Peng, Xionghu Zhong, Hao Huang, Eng Siong Chng

This paper details the experimental results of adapting the OpenAI's Whisper model for Code-Switch Mandarin-English Speech Recognition (ASR) on the SEAME and ASRU2019 corpora. We conducted 2 experiments: a) using adaptation data from 1 to 100/200 hours to demonstrate effectiveness of adaptation, b) examining different language ID setup on Whisper prompt. The Mixed Error Rate results show that the amount of adaptation data may be as low as $1\sim10$ hours to achieve saturation in performance gain (SEAME) while the ASRU task continued to show performance with more adaptation data ($>$100 hours). For the language prompt, the results show that although various prompting strategies initially produce different outcomes, adapting the Whisper model with code-switch data uniformly improves its performance. These results may be relevant also to the community when applying Whisper for related tasks of adapting to new target domains.

📄 PDF Abstract BibTeX arXiv:2311.17382

Code (0)

등록된 구현이 없습니다.

Tasks

speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Leveraging Self-Supervised Models for Automatic Whispered Speech Recognition

2024-07-30 · Aref Farhadipour, Homa Asadi, Volker Dellwo

In automatic speech recognition, any factor that alters the acoustic properties of speech can pose a challenge to the system's performance. This paper presents a novel approach for automatic whispered speech recognition …

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

Swedish Whispers; Leveraging a Massive Speech Corpus for Swedish Speech Recognition

2025-05-23 · Leonora Vesterbacka, Faton Rekathati, Robin Kurtz, Justyna Sikora 외

This work presents a suite of fine-tuned Whisper models for Swedish, trained on a dataset of unprecedented size and variability for this mid-resourced language. As languages of smaller sizes are often underrepresented in…

speech-recognitionSpeech Recognition

Whispering in Norwegian: Navigating Orthographic and Dialectic Challenges

2024-02-02 · Per E Kummervold, Javier de la Rosa, Freddy Wetjen, Rolv-Arild Braaten 외

This article introduces NB-Whisper, an adaptation of OpenAI's Whisper, specifically fine-tuned for Norwegian language Automatic Speech Recognition (ASR). We highlight its key contributions and summarise the results achie…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Calm-Whisper: Reduce Whisper Hallucination On Non-Speech By Calming Crazy Heads Down

2025-05-19 · Yingzhi Wang, Anas Alhmoud, Saad Alsahly, Muhammad Alqurishi 외

OpenAI's Whisper has achieved significant success in Automatic Speech Recognition. However, it has consistently been found to exhibit hallucination issues, particularly in non-speech segments, which limits its broader ap…

Automatic Speech RecognitionDecoderHallucinationspeech-recognition+1

Speech Emotion Recognition Leveraging OpenAI's Whisper Representations and Attentive Pooling Methods

2026-02-05 · Ali Shendabadi, Parnia Izadirad, Mostafa Salehi, Mahmoud Bijankhan arxiv

Speech Emotion Recognition (SER) research has faced limitations due to the lack of standard and sufficiently large datasets. Recent studies have leveraged pre-trained models to extract features for downstream tasks such …

Speech Emotion Recognition