paper-with-me

Papers

BlasBench: An Open Benchmark for Irish Speech Recognition

2026-04-12 · Jyoutir Raj, John Conway arxiv

Existing multilingual benchmarks include Irish among dozens of languages but apply no Irish-aware text normalisation, leaving reliable and reproducible ASR comparison impossible. We introduce BlasBench, an open evaluation harness that provides a standalone Irish-aware normaliser preserving fadas, lenition, and eclipsis; a reproducible scoring harness and per-utterance predictions released for all evaluated runs. We pilot this by benchmarking 12 systems across four architecture families on Common Voice ga-IE and FLEURS ga-IE. All Whisper variants exceed 100% WER through insertion-driven hallucination. Microsoft Azure reaches 22.2% WER on Common Voice and 57.5% on FLEURS; the best open model, Omnilingual ASR 7B, reaches 30.65% and 39.09% respectively. Models fine-tuned on Common Voice degrade 33-43 points moving to FLEURS, while massively multilingual models degrade only 7-10 - a generalisation gap that single-dataset evaluation misses.

📄 PDF Abstract BibTeX arXiv:2604.10736

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Recognition

Similar Papers 제목 키워드 기반

Leveraging Self-Supervised Models for Automatic Whispered Speech Recognition

2024-07-30 · Aref Farhadipour, Homa Asadi, Volker Dellwo

In automatic speech recognition, any factor that alters the acoustic properties of speech can pose a challenge to the system's performance. This paper presents a novel approach for automatic whispered speech recognition …

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

Towards spoken dialect identification of Irish

2023-07-14 · Liam Lonergan, Mengjie Qian, Neasa Ní Chiaráin, Christer Gobl 외

The Irish language is rich in its diversity of dialects and accents. This compounds the difficulty of creating a speech recognition system for the low-resource language, as such a system must contend with a high degree o…

Dialect IdentificationLanguage Identificationspeech-recognitionSpeech Recognition

Low-resource speech recognition and dialect identification of Irish in a multi-task framework

2024-05-02 · Liam Lonergan, Mengjie Qian, Neasa Ní Chiaráin, Christer Gobl 외

This paper explores the use of Hybrid CTC/Attention encoder-decoder models trained with Intermediate CTC (InterCTC) for Irish (Gaelic) low-resource speech recognition (ASR) and dialect identification (DID). Results are c…

DecoderDialect IdentificationLanguage ModelingLanguage Modelling+2

Fotheidil: an Automatic Transcription System for the Irish Language

2024-12-31 · Liam Lonergan, Ibon Saratxaga, John Sloan, Oscar Maharog 외

This paper sets out the first web-based transcription system for the Irish language - Fotheidil, a system that utilises speech-related AI technologies as part of the ABAIR initiative. The system includes both off-the-she…

Action DetectionActivity DetectionAutomatic Speech RecognitionPunctuation Restoration+2

Irish-BLiMP: A Linguistic Benchmark for Evaluating Human and Language Model Performance in a Low-Resource Setting

2025-10-23 · Josh McGiff, Khanh-Tung Tran, William Mulcahy, Dáibhidh Ó Luinín 외 arxiv

We present Irish-BLiMP (Irish Benchmark of Linguistic Minimal Pairs), the first dataset and framework designed for fine-grained evaluation of linguistic competence in the Irish language, an endangered language. Drawing o…