paper-with-me

Papers

Free English and Czech telephone speech corpus shared under the CC-BY-SA 3.0 license

2014-05-01 · LREC 2014 5 · Mat{\v{e}}j Korvas, Ond{\v{r}}ej Pl{\'a}tek, Ond{\v{r}}ej Du{\v{s}}ek, Luk{\'a}{\v{s}} {\v{Z}}ilka, Filip Jur{\v{c}}{\'\i}{\v{c}}ek

We present a dataset of telephone conversations in English and Czech, developed for training acoustic models for automatic speech recognition (ASR) in spoken dialogue systems (SDSs). The data comprise 45 hours of speech in English and over 18 hours in Czech. Large part of the data, both audio and transcriptions, was collected using crowdsourcing, the rest are transcriptions by hired transcribers. We release the data together with scripts for data pre-processing and building acoustic models using the HTK and Kaldi ASR toolkits. We publish also the trained models described in this paper. The data are released under the CC-BY-SA{\textasciitilde}3.0 license, the scripts are licensed under Apache{\textasciitilde}2.0. In the paper, we report on the methodology of collecting the data, on the size and properties of the data, and on the scripts and their use. We verify the usability of the datasets by training and evaluating acoustic models using the presented data and scripts.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Domain Adaptationspeech-recognitionSpeech RecognitionSpoken Dialogue Systems

Similar Papers 제목 키워드 기반

Advancing Speech Translation: A Corpus of Mandarin-English Conversational Telephone Speech

2024-03-25 · Shannon Wotherspoon, William Hartmann, Matthew Snover

This paper introduces a set of English translations for a 123-hour subset of the CallHome Mandarin Chinese data and the HKUST Mandarin Telephone Speech data for the task of speech translation. Paired source-language spee…

Translation

The Nijmegen Corpus of Casual Czech

2014-05-01 · LREC 2014 5 · Mirjam Ernestus, Lucie Ko{\v{c}}kov{\'a}-Amortov{\'a}, Petr Pollak

This article introduces a new speech corpus, the Nijmegen Corpus of Casual Czech (NCCCz), which contains more than 30 hours of high-quality recordings of casual conversations in Common Czech, among ten groups of three ma…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

AI Brown and AI Koditex: LLM-Generated Corpora Comparable to Traditional Corpora of English and Czech Texts

2025-09-26 · Jiří Milička, Anna Marklová, Václav Cvrček arxiv

This article presents two corpora of English and Czech texts generated with large language models (LLMs). The motivation is to create a resource for comparing human-written texts with LLM-generated text linguistically. E…

The Marchex 2018 English Conversational Telephone Speech Recognition System

2018-11-05 · Seongjun Hahm, Iroro Orife, Shane Walker, Jason Flaks

In this paper, we describe recent performance improvements to the production Marchex speech recognition system for our spontaneous customer-to-business telephone conversations. In our previous work, we focused on in-doma…

Language ModelingLanguage Modellingspeech-recognitionSpeech Recognition

WeCanTalk: A New Multi-language, Multi-modal Resource for Speaker Recognition

2022-06-01 · LREC 2022 6 · Karen Jones, Kevin Walker, Christopher Caruso, Jonathan Wright 외

The WeCanTalk (WCT) Corpus is a new multi-language, multi-modal resource for speaker recognition. The corpus contains Cantonese, Mandarin and English telephony and video speech data from over 200 multilingual speakers lo…

Speaker Recognition