paper-with-me

홈 › Papers

BaltiVoice: A Speech Corpus and Fine-tuned Whisper ASR System for the Balti Language

2026-06-02 · Muhammad Ali arxiv

We present BaltiVoice, a 16.8-hour read-speech corpus for Balti (ISO 639-3: bft), a Tibetic language spoken in Gilgit-Baltistan, Pakistan, with no prior publicly available ASR resources. The corpus contains 10,060 validated utterances in native Nastaliq script, derived from Mozilla Common Voice recordings. Fine-tuning OpenAI Whisper-small yields a Word Error Rate (WER) of 24.78% and a Character Error Rate (CER) of 8.30% after training for 5 epochs (3,000 steps) on the 538-utterance speaker-disjoint validation set, down from a zero-shot baseline of 159.19% WER and 152.52% CER. A Whisper-base fine-tuned on the same data achieves 44.54% WER and 15.61% CER, confirming that model capacity matters for this low-resource setting. The dataset, fine-tuned model, and a live transcription demo are publicly available on HuggingFace.

📄 PDF Abstract BibTeX arXiv:2606.03504

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Improving the Inclusivity of Dutch Speech Recognition by Fine-tuning Whisper on the JASMIN-CGN Corpus

2025-02-24 · Golshid Shekoufandeh, Paul Boersma, Antal Van den Bosch

We test and study the variation in speech recognition of fine-tuned versions of the Whisper model on child, elderly and non-native Dutch speech from the JASMIN-CGN corpus. Our primary goal is to evaluate how speakers' ag…

Automatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

SloPal: A 60-Million-Word Slovak Parliamentary Corpus with Aligned Speech and Fine-Tuned ASR Models

2025-09-23 · Erik Božík, Marek Šuppa arxiv

Slovak remains a low-resource language for automatic speech recognition (ASR), with fewer than 100 hours of publicly available training data. We present SloPal, a comprehensive Slovak parliamentary corpus comprising 330,…

Speech Recognition

Swedish Whispers; Leveraging a Massive Speech Corpus for Swedish Speech Recognition

2025-05-23 · Leonora Vesterbacka, Faton Rekathati, Robin Kurtz, Justyna Sikora 외

This work presents a suite of fine-tuned Whisper models for Swedish, trained on a dataset of unprecedented size and variability for this mid-resourced language. As languages of smaller sizes are often underrepresented in…

speech-recognitionSpeech Recognition

Adaptation of Whisper models to child speech recognition

2023-07-24 · Rishabh Jain, Andrei Barcovschi, Mariam Yiwere, Peter Corcoran 외

Automatic Speech Recognition (ASR) systems often struggle with transcribing child speech due to the lack of large child speech datasets required to accurately train child-friendly ASR models. However, there are huge amou…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

BEST-RQ-Based Self-Supervised Learning for Whisper Domain Adaptation

2025-10-28 · Raphaël Bagat, Irina Illina, Emmanuel Vincent arxiv

Automatic Speech Recognition (ASR) systems, despite large multilingual training, struggle in low-resource scenarios where labeled data is scarce. We propose BEARD (BEST-RQ Encoder Adaptation with Re-training and Distilla…

Self-Supervised LearningKnowledge DistillationSpeech RecognitionDomain Adaptation