paper-with-me

Papers

Analysis of Data Augmentation Methods for Low-Resource Maltese ASR

2021-11-15 · Andrea DeMarco, Carlos Mena, Albert Gatt, Claudia Borg, Aiden Williams, Lonneke van der Plas

Recent years have seen an increased interest in the computational speech processing of Maltese, but resources remain sparse. In this paper, we consider data augmentation techniques for improving speech recognition for low-resource languages, focusing on Maltese as a test case. We consider three different types of data augmentation: unsupervised training, multilingual training and the use of synthesized speech as training data. The goal is to determine which of these techniques, or combination of them, is the most effective to improve speech recognition for languages where the starting point is a small corpus of approximately 7 hours of transcribed speech. Our results show that combining the data augmentation techniques studied here lead us to an absolute WER improvement of 15% without the use of a language model.

📄 PDF Abstract BibTeX arXiv:2111.07793

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationLanguage ModelingLanguage Modellingspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Data Augmentation for Maltese NLP using Transliterated and Machine Translated Arabic Data

2025-09-16 · Kurt Micallef, Nizar Habash, Claudia Borg arxiv

Maltese is a unique Semitic language that has evolved under extensive influence from Romance and Germanic languages, particularly Italian and English. Despite its Semitic roots, its orthography is based on the Latin scri…

Machine TranslationData Augmentation

From Measurement to Mitigation: Exploring the Transferability of Debiasing Approaches to Gender Bias in Maltese Language Models

2025-07-03 · Melanie Galea, Claudia Borg arxiv

The advancement of Large Language Models (LLMs) has transformed Natural Language Processing (NLP), enabling performance across diverse tasks with little task-specific training. However, LLMs remain susceptible to social …

Data Augmentation

Baseline English and Maltese-English Classification Models for Subjectivity Detection, Sentiment Analysis, Emotion Analysis, Sarcasm Detection, and Irony Detection

2022-06-01 · SIGUL (LREC) 2022 6 · Keith Cortis, Brian Davis

This paper presents baseline classification models for subjectivity detection, sentiment analysis, emotion analysis, sarcasm detection, and irony detection. All models are trained on user-generated content gathered from …

ClassificationEmotion RecognitionregressionSarcasm Detection+1

Fine-tuning Neural Language Models for Multidimensional Opinion Mining of English-Maltese Social Data

2021-09-01 · RANLP 2021 9 · Keith Cortis, Kanishk Verma, Brian Davis

This paper presents multidimensional Social Opinion Mining on user-generated content gathered from newswires and social networking services in three different languages: English —a high-resourced language, Maltese —a low…

ClassificationOpinion Mining

Malta National Language Technology Platform: A vision for enhancing Malta’s official languages using Machine Translation

2021-09-01 · MMTLRL (RANLP) 2021 9 · Keith Cortis, Judie Attard, Donatienne Spiteri

In this paper we introduce a vision towards establishing the Malta National Language Technology Platform; an ongoing effort that aims to provide a basis for enhancing Malta’s official languages, namely Maltese and Englis…

Machine TranslationTranslation