paper-with-me

홈 › Papers

The SI TEDx-UM speech database: a new Slovenian Spoken Language Resource

2016-05-01 · LREC 2016 5 · Andrej {\v{Z}}gank, Mirjam Sepesy Mau{\v{c}}ec, Darinka Verdonik

This paper presents a new Slovenian spoken language resource built from TEDx Talks. The speech database contains 242 talks in total duration of 54 hours. The annotation and transcription of acquired spoken material was generated automatically, applying acoustic segmentation and automatic speech recognition. The development and evaluation subset was also manually transcribed using the guidelines specified for the Slovenian GOS corpus. The manual transcriptions were used to evaluate the quality of unsupervised transcriptions. The average word error rate for the SI TEDx-UM evaluation subset was 50.7{\%}, with out of vocabulary rate of 24{\%} and language model perplexity of 390. The unsupervised transcriptions contain 372k tokens, where 32k of them were different.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modellingspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

The Universal Dependencies Treebank of Spoken Slovenian

2016-05-01 · LREC 2016 5 · Kaja Dobrovoljc, Joakim Nivre

This paper presents the construction of an open-source dependency treebank of spoken Slovenian, the first syntactically annotated collection of spontaneous speech in Slovenian. The treebank has been manually annotated us…

Annotating formulaic sequences in spoken Slovenian: structure, function and relevance

2019-08-01 · WS 2019 8 · Kaja Dobrovoljc

This paper presents the identification of formulaic sequences in the reference corpus of spoken Slovenian and their annotation in terms of syntactic structure, pragmatic function and lexicographic relevance. The annotati…

The Multilingual TEDx Corpus for Speech Recognition and Translation

2021-02-02 · Elizabeth Salesky, Matthew Wiesner, Jacob Bremerman, Roldano Cattoni 외

We present the Multilingual TEDx corpus, built to support speech recognition (ASR) and speech translation (ST) research across many non-English source languages. The corpus is a collection of audio recordings from TEDx t…

speech-recognitionSpeech RecognitionTranslation

TEDxTN: A Three-way Speech Translation Corpus for Code-Switched Tunisian Arabic - English

2025-11-13 · Fethi Bougares, Salima Mdhaffar, Haroun Elleuch, Yannick Estève arxiv

In this paper, we introduce TEDxTN, the first publicly available Tunisian Arabic to English speech translation dataset. This work is in line with the ongoing effort to mitigate the data scarcity obstacle for a number of …

Speech Recognition

Er ... well, it matters, right? On the role of data representations in spoken language dependency parsing

2018-11-01 · WS 2018 11 · Kaja Dobrovoljc, Matej Martinc

Despite the significant improvement of data-driven dependency parsing systems in recent years, they still achieve a considerably lower performance in parsing spoken language data in comparison to written data. On the exa…

Dependency ParsingLanguage Modelling