paper-with-me

Papers

Exploring Capabilities of Monolingual Audio Transformers using Large Datasets in Automatic Speech Recognition of Czech

2022-06-15 · Jan Lehečka, Jan Švec, Aleš Pražák, Josef V. Psutka

In this paper, we present our progress in pretraining Czech monolingual audio transformers from a large dataset containing more than 80 thousand hours of unlabeled speech, and subsequently fine-tuning the model on automatic speech recognition tasks using a combination of in-domain data and almost 6 thousand hours of out-of-domain transcribed speech. We are presenting a large palette of experiments with various fine-tuning setups evaluated on two public datasets (CommonVoice and VoxPopuli) and one extremely challenging dataset from the MALACH project. Our results show that monolingual Wav2Vec 2.0 models are robust ASR systems, which can take advantage of large labeled and unlabeled datasets and successfully compete with state-of-the-art LVCSR systems. Moreover, Wav2Vec models proved to be good zero-shot learners when no training data are available for the target ASR task.

📄 PDF Abstract BibTeX arXiv:2206.07627

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Exploring Linguistic Properties of Monolingual BERTs with Typological Classification among Languages

2023-05-03 · Elena Sofia Ruzzetti, Federico Ranaldi, Felicia Logozzo, Michele Mastromattei 외

The impressive achievements of transformers force NLP researchers to delve into how these models represent the underlying structure of natural language. In this paper, we propose a novel standpoint to investigate the abo…

Domain Adaptation

Prompting Large Language Models with Speech Recognition Abilities

2023-07-21 · Yassir Fathullah, Chunyang Wu, Egor Lakomkin, Junteng Jia 외

Large language models have proven themselves highly flexible, able to solve a wide range of generative tasks, such as abstractive summarization and open-ended question answering. In this paper we extend the capabilities …

Abstractive Text SummarizationAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Open-Ended Question Answering+3

Scaling Auditory Cognition via Test-Time Compute in Audio Language Models

2025-03-30 · Ting Dang, Yan Gao, Hong Jia

Large language models (LLMs) have shown exceptional versatility in natural language processing, prompting recent efforts to extend their multimodal capabilities to speech processing through the development of audio large…

speech-recognitionSpeech Recognition

Optimizing ASR for Catalan-Spanish Code-Switching: A Comparative Analysis of Methodologies

2025-07-18 · Carlos Mena, Pol Serra, Jacobo Romero, Abir Messaoudi 외 arxiv

Code-switching (CS), the alternating use of two or more languages, challenges automatic speech recognition (ASR) due to scarce training data and linguistic similarities. The lack of dedicated CS datasets limits ASR perfo…

Speech Recognition

Exploring Spoken Language Identification Strategies for Automatic Transcription of Multilingual Broadcast and Institutional Speech

2024-06-13 · Martina Valente, Fabio Brugnara, Giovanni Morrone, Enrico Zovato 외

This paper addresses spoken language identification (SLI) and speech recognition of multilingual broadcast and institutional speech, real application scenarios that have been rarely addressed in the SLI literature. Obser…

Language Identificationspeaker-diarizationSpeaker Diarizationspeech-recognition+2