paper-with-me

홈 › Papers

KALAKA-2: a TV Broadcast Speech Database for the Recognition of Iberian Languages in Clean and Noisy Environments

2012-05-01 · LREC 2012 5 · Luis Javier Rodr{\'\i}guez-Fuentes, Mikel Penagarikano, Amparo Varona, Mireia Diez, Germ{\'a}n Bordel

This paper presents the main features (design issues, recording setup, etc.) of KALAKA-2, a TV broadcast speech database specifically designed for the development and evaluation of language recognition systems in clean and noisy environments. KALAKA-2 was created to support the Albayzin 2010 Language Recognition Evaluation (LRE), organized by the Spanish Network on Speech Technologies from June to November 2010. The database features 6 target languages: Basque, Catalan, English, Galician, Portuguese and Spanish, and includes segments in other (Out-Of-Set) languages, which allow to perform open-set verification tests. The best performance attained in the Albayzin 2010 LRE is presented and briefly discussed. The performance of a state-of-the-art system in various tasks defined on the database is also presented. In both cases, results highlight the suitability of KALAKA-2 as a benchmark for the development and evaluation of language recognition technology.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

KALAKA-3: a database for the recognition of spoken European languages on YouTube audios

2014-05-01 · LREC 2014 5 · Luis Javier Rodr{\'\i}guez-Fuentes, Mikel Penagarikano, Amparo Varona, Mireia Diez 외

This paper describes the main features of KALAKA-3, a speech database specifically designed for the development and evaluation of language recognition systems. The database provides TV broadcast speech for training, and …

Development of the Siberian Ingrian Finnish Speech Corpus

2022-05-01 · ComputEL (ACL) 2022 5 · Ivan Ubaleht, Taisto-Kalevi Raudalainen

In this paper we present the speech corpus for the Siberian Ingrian Finnish language. The speech corpus includes audio data, annotations, software tools for data-processing, two databases and a web application. We have p…

The Slovene BNSI Broadcast News database and reference speech corpus GOS: Towards the uniform guidelines for future work

2014-05-01 · LREC 2014 5 · Andrej {\v{Z}}gank, Ana Zwitter Vitez, Darinka Verdonik

The aim of the paper is to search for common guidelines for the future development of speech databases for less resourced languages in order to make them the most useful for both main fields of their use, linguistic rese…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

TUKE-BNews-SK: Slovak Broadcast News Corpus Construction and Evaluation

2014-05-01 · LREC 2014 5 · Mat{\'u}{\v{s}} Pleva, Jozef Juh{\'a}r

This article presents an overview of the existing acoustical corpuses suitable for broadcast news automatic transcription task in the Slovak language. The TUKE-BNews-SK database created in our department was built to sup…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

New bilingual speech databases for audio diarization

2014-05-01 · LREC 2014 5 · David Tavarez, Eva Navas, Daniel Erro, Ibon Saratxaga 외

This paper describes the process of collecting and recording two new bilingual speech databases in Spanish and Basque. They are designed primarily for speaker diarization in two different application domains: broadcast n…

speaker-diarizationSpeaker DiarizationSpeaker Recognition