Creation of a Balanced State-of-the-Art Multilayer Corpus for NLU
Code (1)
Tasks
Abstractive Text SummarizationCoreference ResolutionEntity LinkingKnowledge Base PopulationNamed Entity Recognition (NER)Semantic ParsingSemantic Role LabelingText SummarizationSimilar Papers 제목 키워드 기반
Designing the Latvian Speech Recognition Corpus
In this paper the authors present the first Latvian speech corpus designed specifically for speech recognition purposes. The paper outlines the decisions made in the corpus designing process through analysis of related w…
speech-recognitionSpeech RecognitionSpeech SynthesisAMALGUM -- A Free, Balanced, Multilayer English Web Corpus
We present a freely available, genre-balanced English web corpus totaling 4M tokens and featuring a large number of high-quality automatic annotation layers, including dependency trees, non-named entity annotations, core…
coreference-resolutionCoreference ResolutionOpera Graeca Adnotata: Building a 34M+ Token Multilayer Corpus for Ancient Greek
In this article, the beta version 0.1.0 of Opera Graeca Adnotata (OGA), the largest open-access multilayer corpus for Ancient Greek (AG) is presented. OGA consists of 1,687 literary works and 34M+ tokens coming from the …
LemmatizationSentenceSentence segmentationDesign of a Tigrinya Language Speech Corpus for Speech Recognition
In this paper, we describe the first Tigrinya Languages speech corpora designed and development for speech recognition purposes. Tigrinya, often written as Tigrigna (ትግርኛ) /tɪˈɡrinjə/ belongs to the Semitic branch of the…
speech-recognitionSpeech RecognitionAudiobook Dialogues as Training Data for Conversational Style Synthetic Voices
Synthetic voices are increasingly used in applications that require a conversational speaking style, raising the question as to which type of training data yields the most suitable speaking style for such applications. T…
Sentencetext-to-speechText to Speech