BulPhonC: Bulgarian Speech Corpus for the Development of ASR Technology
In this paper we introduce a Bulgarian speech database, which was created for the purpose of ASR technology development. The paper describes the design and the content of the speech database. We present also an empirical evaluation of the performance of a LVCSR system for Bulgarian trained on the BulPhonC data. The resource is available free for scientific usage.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
The Political Speech Corpus of Bulgarian
The paper introduces the Political Speech Corpus of Bulgarian. First, its current state has been discussed with respect to its size, coverage, genre specification and related online services. Then, the focus goes to the …
LemmatizationMorphological AnalysisSentiment AnalysisSyntactic characteristics of emotive predicates in Bulgarian: A corpus-based study
The paper presents a corpus-based study of emotive predicates (verbs and predicative constructions with adjectival, adverbial or noun phrases) in Bulgarian with respect to their syntactic characteristics. The sources of …
Feature-Rich Part-of-speech Tagging for Morphologically Complex Languages: Application to Bulgarian
We present experiments with part-of-speech tagging for Bulgarian, a Slavic language with rich inflectional and derivational morphology. Unlike most previous work, which has used a small number of grammatical categories, …
Part-Of-Speech TaggingPOSAnnotation of Clinical Narratives in Bulgarian language
In this paper we describe annotation process of clinical texts with morphosyntactic and semantic information. The corpus contains 1,300 discharge letters in Bulgarian language for patients with Endocrinology and Metaboli…
ChunkingDependency ParsingInformation RetrievalCategorisation of Bulgarian Legislative Documents
The paper presents the categorisation of Bulgarian MARCELL corpus in toplevel EuroVoc domains. The Bulgarian MARCELL corpus is part of a recently developed multilingual corpus representing the national legislation in sev…
Term Extraction