paper-with-me

Papers

A Comparative Analysis of Bilingual and Trilingual Wav2Vec Models for Automatic Speech Recognition in Multilingual Oral History Archives

2024-07-24 · Jan Lehečka, Josef V. Psutka, Luboš Šmídl, Pavel Ircing, Josef Psutka

In this paper, we are comparing monolingual Wav2Vec 2.0 models with various multilingual models to see whether we could improve speech recognition performance on a unique oral history archive containing a lot of mixed-language sentences. Our main goal is to push forward research on this unique dataset, which is an extremely valuable part of our cultural heritage. Our results suggest that monolingual speech recognition models are, in most cases, superior to multilingual models, even when processing the oral history archive full of mixed-language sentences from non-native speakers. We also performed the same experiments on the public CommonVoice dataset to verify our results. We are contributing to the research community by releasing our pre-trained models to the public.

📄 PDF Abstract BibTeX arXiv:2407.17160

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

The Trilingual ALLEGRA Corpus: Presentation and Possible Use for Lexicon Induction

2012-05-01 · LREC 2012 5 · Yves Scherrer, Bruno Cartoni

In this paper, we present a trilingual parallel corpus for German, Italian and Romansh, a Swiss minority language spoken in the canton of Grisons. The corpus called ALLEGRA contains press releases automatically gathered …

Sentence

Advancing Multilingual Pre-training: TRIP Triangular Document-level Pre-training for Multilingual Language Models

2022-12-15 · Hongyuan Lu, Haoyang Huang, Shuming Ma, Dongdong Zhang 외

Despite the success of multilingual sequence-to-sequence pre-training, most existing approaches rely on document-level monolingual corpora in many different languages, sentence-level bilingual corpora,\footnote{In this p…

Abstractive Text SummarizationCross-Lingual Abstractive SummarizationDocument Level Machine TranslationMachine Translation+2

BBPE16: UTF-16-based byte-level byte-pair encoding for improved multilingual speech recognition

2026-02-02 · Hyunsik Kim, Haeri Kim, Munhak Lee, Kyungmin Lee arxiv

Multilingual automatic speech recognition (ASR) requires tokenization that efficiently covers many writing systems. Byte-level BPE (BBPE) using UTF-8 is widely adopted for its language-agnostic design and full Unicode co…

Speech Recognition

The I3MEDIA speech database: a trilingual annotated corpus for the analysis and synthesis of emotional speech

2012-05-01 · LREC 2012 5 · Juan Mar{\'\i}a Garrido, Yesika Laplaza, Montse Marquina, Andrea Pearman 외

In this article the I3Media corpus is presented, a trilingual (Catalan, English, Spanish) speech database of neutral and emotional material collected for analysis and synthesis purposes. The corpus is actually made up of…

MultiMed-ST: Large-scale Many-to-many Multilingual Medical Speech Translation

2025-04-04 · Khai Le-Duc, Tuyen Tran, Bach Phan Tat, Nguyen Kim Hai Bui 외

Multilingual speech translation (ST) in the medical domain enhances patient care by enabling efficient communication across language barriers, alleviating specialized workforce shortages, and facilitating improved diagno…

Machine TranslationTranslation