paper-with-me

Papers

MASRI-HEADSET: A Maltese Corpus for Speech Recognition

2020-08-13 · LREC 2020 5 · Carlos Mena, Albert Gatt, Andrea DeMarco, Claudia Borg, Lonneke van der Plas, Amanda Muscat, Ian Padovani

Maltese, the national language of Malta, is spoken by approximately 500,000 people. Speech processing for Maltese is still in its early stages of development. In this paper, we present the first spoken Maltese corpus designed purposely for Automatic Speech Recognition (ASR). The MASRI-HEADSET corpus was developed by the MASRI project at the University of Malta. It consists of 8 hours of speech paired with text, recorded by using short text snippets in a laboratory environment. The speakers were recruited from different geographical locations all over the Maltese islands, and were roughly evenly distributed by gender. This paper also presents some initial results achieved in baseline experiments for Maltese ASR using Sphinx and Kaldi. The MASRI-HEADSET Corpus is publicly available for research/academic purposes.

📄 PDF Abstract BibTeX arXiv:2008.05760

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Analysis of Data Augmentation Methods for Low-Resource Maltese ASR

2021-11-15 · Andrea DeMarco, Carlos Mena, Albert Gatt, Claudia Borg 외

Recent years have seen an increased interest in the computational speech processing of Maltese, but resources remain sparse. In this paper, we consider data augmentation techniques for improving speech recognition for lo…

Data AugmentationLanguage ModelingLanguage Modellingspeech-recognition+1

Pre-training Data Quality and Quantity for a Low-Resource Language: New Corpus and BERT Models for Maltese

2022-05-21 · DeepLo 2022 7 · Kurt Micallef, Albert Gatt, Marc Tanti, Lonneke van der Plas 외

Multilingual language models such as mBERT have seen impressive cross-lingual transfer to a variety of languages, but many languages remain excluded from these models. In this paper, we analyse the effect of pre-training…

Cross-Lingual TransferDependency ParsingNamed Entity RecognitionNamed Entity Recognition (NER)+2

LV-ROVER-MLT: Low-Resource Maltese OCR by Synthetic Fine-Tuning and Multi-Stream Arbitration

2026-06-30 · Adam Darmanin arxiv

Maltese OCR is constrained by the absence of a public, reusable paragraph-scale training corpus. We address this by generating synthetic Maltese line images, fine-tuning the Tesseract 5 LSTM, and combining five determini…

Incorporating an Error Corpus into a Spellchecker for Maltese

2012-05-01 · LREC 2012 5 · Michael Rosner, Albert Gatt, Andrew Attard, Jan Joachimsen

This paper discusses the ongoing development of a new Maltese spell checker, highlighting the methodologies which would best suit such a language. We thus discuss several previous attempts, highlighting what we believe t…

Malta National Language Technology Platform: A vision for enhancing Malta’s official languages using Machine Translation

2021-09-01 · MMTLRL (RANLP) 2021 9 · Keith Cortis, Judie Attard, Donatienne Spiteri

In this paper we introduce a vision towards establishing the Malta National Language Technology Platform; an ongoing effort that aims to provide a basis for enhancing Malta’s official languages, namely Maltese and Englis…

Machine TranslationTranslation