paper-with-me

홈 › Papers

Towards dialect-inclusive recognition in a low-resource language: are balanced corpora the answer?

2023-07-14 · Liam Lonergan, Mengjie Qian, Neasa Ní Chiaráin, Christer Gobl, Ailbhe Ní Chasaide

ASR systems are generally built for the spoken 'standard', and their performance declines for non-standard dialects/varieties. This is a problem for a language like Irish, where there is no single spoken standard, but rather three major dialects: Ulster (Ul), Connacht (Co) and Munster (Mu). As a diagnostic to quantify the effect of the speaker's dialect on recognition performance, 12 ASR systems were trained, firstly using baseline dialect-balanced training corpora, and then using modified versions of the baseline corpora, where dialect-specific materials were either subtracted or added. Results indicate that dialect-balanced corpora do not yield a similar performance across the dialects: the Ul dialect consistently underperforms, whereas Mu yields lowest WERs. There is a close relationship between Co and Mu dialects, but one that is not symmetrical. These results will guide future corpus collection and system building strategies to optimise for cross-dialect performance equity.

📄 PDF Abstract BibTeX arXiv:2307.07295

Code (0)

등록된 구현이 없습니다.

Tasks

Diagnostic

Similar Papers 제목 키워드 기반

BIDWESH: A Bangla Regional Based Hate Speech Detection Dataset

2025-07-22 · Azizul Hakim Fayaz, MD. Shorif Uddin, Rayhan Uddin Bhuiyan, Zakia Sultana 외 arxiv

Hate speech on digital platforms has become a growing concern globally, especially in linguistically diverse countries like Bangladesh, where regional dialects play a major role in everyday communication. Despite progres…

Hate Speech Detection

Towards spoken dialect identification of Irish

2023-07-14 · Liam Lonergan, Mengjie Qian, Neasa Ní Chiaráin, Christer Gobl 외

The Irish language is rich in its diversity of dialects and accents. This compounds the difficulty of creating a speech recognition system for the low-resource language, as such a system must contend with a high degree o…

Dialect IdentificationLanguage Identificationspeech-recognitionSpeech Recognition

BanglaDialecto: An End-to-End AI-Powered Regional Speech Standardization

2024-11-16 · Md. Nazmus Sadat Samin, Jawad Ibn Ahad, Tanjila Ahmed Medha, Fuad Rahman 외

This study focuses on recognizing Bangladeshi dialects and converting diverse Bengali accents into standardized formal Bengali speech. Dialects, often referred to as regional languages, are distinctive variations of a la…

Machine Translationspeech-recognitionSpeech Recognition

Multi-VALUE: A Framework for Cross-Dialectal English NLP

2022-12-15 · Caleb Ziems, William Held, Jingfeng Yang, Jwala Dhamala 외

Dialect differences caused by regional, social, and economic factors cause performance discrepancies for many groups of language technology users. Inclusive and equitable language technology must critically be dialect in…

Data AugmentationMachine TranslationQuestion AnsweringSemantic Parsing+1

MoMQ: Mixture-of-Experts Enhances Multi-Dialect Query Generation across Relational and Non-Relational Databases

2024-10-24 · Zhisheng Lin, Yifu Liu, Zhiling Luo, Jinyang Gao 외

The improvement in translating natural language to structured query language (SQL) can be attributed to the advancements in large language models (LLMs). Open-source LLMs, tailored for specific database dialects such as …

Mixture-of-Experts