paper-with-me

Papers

Using Kaldi for Automatic Speech Recognition of Conversational Austrian German

2023-01-16 · Julian Linke, Saskia Wepner, Gernot Kubin, Barbara Schuppler

As dialogue systems are becoming more and more interactional and social, also the accurate automatic speech recognition (ASR) of conversational speech is of increasing importance. This shifts the focus from short, spontaneous, task-oriented dialogues to the much higher complexity of casual face-to-face conversations. However, the collection and annotation of such conversations is a time-consuming process and data is sparse for this specific speaking style. This paper presents ASR experiments with read and conversational Austrian German as target. In order to deal with having only limited resources available for conversational German and, at the same time, with a large variation among speakers with respect to pronunciation characteristics, we improve a Kaldi-based ASR system by incorporating a (large) knowledge-based pronunciation lexicon, while exploring different data-based methods to restrict the number of pronunciation variants for each lexical entry. We achieve best WER of 0.4% on Austrian German read speech and best average WER of 48.5% on conversational speech. We find that by using our best pronunciation lexicon a similarly high performance can be achieved than by increasing the size of the data used for the language model by approx. 360% to 760%. Our findings indicate that for low-resource scenarios -- despite the general trend in speech technology towards using data-based methods only -- knowledge-based approaches are a successful, efficient method.

📄 PDF Abstract BibTeX arXiv:2301.06475

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language Modellingspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Conversational Speech Recognition Needs Data? Experiments with Austrian German

2022-06-01 · LREC 2022 6 · Julian Linke, Philip N. Garner, Gernot Kubin, Barbara Schuppler

Conversational speech represents one of the most complex of automatic speech recognition (ASR) tasks owing to the high inter-speaker variation in both pronunciation and conversational dynamics. Such complexity is particu…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+2

ExKaldi-RT: A Real-Time Automatic Speech Recognition Extension Toolkit of Kaldi

2021-04-03 · Yu Wang, Chee Siang Leow, Akio Kobayashi, Takehito Utsuro 외

This paper describes the ExKaldi-RT online automatic speech recognition (ASR) toolkit that is implemented based on the Kaldi ASR toolkit and Python language. ExKaldi-RT provides tools for building online recognition pipe…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Prominence-aware automatic speech recognition for conversational speech

2025-09-12 · Julian Linke, Barbara Schuppler arxiv

This paper investigates prominence-aware automatic speech recognition (ASR) by combining prominence detection and speech recognition for conversational Austrian German. First, prominence detectors were developed by fine-…

Speech Recognition

Automated speech audiometry: Can it work using open-source pre-trained Kaldi-NL automatic speech recognition?

2023-12-19 · Gloria Araiza-Illan, Luke Meyer, Khiet P. Truong, Deniz Baskent

A practical speech audiometry tool is the digits-in-noise (DIN) test for hearing screening of populations of varying ages and hearing status. The test is usually conducted by a human supervisor (e.g., clinician), who sco…

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

A non-expert Kaldi recipe for Vietnamese Speech Recognition System

2016-12-01 · WS 2016 12 · Hieu-Thi Luong, Hai-Quan Vu

In this paper we describe a non-expert setup for Vietnamese speech recognition system using Kaldi toolkit. We collected a speech corpus over fifteen hours from about fifty Vietnamese native speakers and using it to test …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1