paper-with-me

Papers

Framework for Curating Speech Datasets and Evaluating ASR Systems: A Case Study for Polish

2024-07-18 · Michał Junczyk

Speech datasets available in the public domain are often underutilized because of challenges in discoverability and interoperability. A comprehensive framework has been designed to survey, catalog, and curate available speech datasets, which allows replicable evaluation of automatic speech recognition (ASR) systems. A case study focused on the Polish language was conducted; the framework was applied to curate more than 24 datasets and evaluate 25 combinations of ASR systems and models. This research constitutes the most extensive comparison to date of both commercial and free ASR systems for the Polish language. It draws insights from 600 system-model-test set evaluations, marking a significant advancement in both scale and comprehensiveness. The results of surveys and performance comparisons are available as interactive dashboards (https://huggingface.co/spaces/amu-cai/pl-asr-leaderboard) along with curated datasets (https://huggingface.co/datasets/amu-cai/pl-asr-bigos-v2, https://huggingface.co/datasets/pelcra/pl-asr-pelcra-for-bigos) and the open challenge call (https://poleval.pl/tasks/task3). Tools used for evaluation are open-sourced (https://github.com/goodmike31/pl-asr-bigos-tools), facilitating replication and adaptation for other languages, as well as continuous expansion with new datasets and systems.

📄 PDF Abstract BibTeX arXiv:2408.00005

Code (1)

goodmike31/pl-asr-bigos-tools 공식 구현

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Towards Robust Speech Recognition for Jamaican Patois Music Transcription

2025-07-15 · Jordan Madden, Matthew Stone, Dimitri Johnson, Daniel Geddez arxiv

Although Jamaican Patois is a widely spoken language, current speech recognition systems perform poorly on Patois music, producing inaccurate captions that limit accessibility and hinder downstream applications. In this …

Music TranscriptionSpeech Recognition

TTSDS -- Text-to-Speech Distribution Score

2024-07-17 · Christoph Minixhofer, Ondřej Klejch, Peter Bell

Many recently published Text-to-Speech (TTS) systems produce audio close to real speech. However, TTS evaluation needs to be revisited to make sense of the results obtained with the new architectures, approaches and data…

text-to-speechText to Speech

ESB: A Benchmark For Multi-Domain End-to-End Speech Recognition

2022-10-24 · Sanchit Gandhi, Patrick von Platen, Alexander M. Rush

Speech recognition applications cover a range of different audio and text distributions, with different speaking styles, background noise, transcription punctuation and character casing. However, many speech recognition …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)BenchmarkingSpeech Recognition

Separating Hate Speech and Offensive Language Classes via Adversarial Debiasing

2022-07-01 · NAACL (WOAH) 2022 7 · Shuzhou Yuan, Antonis Maronikolakis, Hinrich Schütze

Research to tackle hate speech plaguing online media has made strides in providing solutions, analyzing bias and curating data. A challenging problem is ambiguity between hate speech and offensive language, causing low p…

SpeechLMScore: Evaluating speech generation using speech language model

2022-12-08 · Soumi Maiti, Yifan Peng, Takaaki Saeki, Shinji Watanabe

While human evaluation is the most reliable metric for evaluating speech generation systems, it is generally costly and time-consuming. Previous studies on automatic speech quality assessment address the problem by predi…

Language ModelingLanguage ModellingmodelSpeech Enhancement+3