Speech Segmentation Optimization using Segmented Bilingual Speech Corpus for End-to-end Speech Translation
Speech segmentation, which splits long speech into short segments, is essential for speech translation (ST). Popular VAD tools like WebRTC VAD have generally relied on pause-based segmentation. Unfortunately, pauses in speech do not necessarily match sentence boundaries, and sentences can be connected by a very short pause that is difficult to detect by VAD. In this study, we propose a speech segmentation method using a binary classification model trained using a segmented bilingual speech corpus. We also propose a hybrid method that combines VAD and the above speech segmentation method. Experimental results revealed that the proposed method is more suitable for cascade and end-to-end ST systems than conventional segmentation methods. The hybrid approach further improved the translation performance.
Code (1)
Tasks
Binary ClassificationSegmentationSentenceTranslationSimilar Papers 제목 키워드 기반
Unsupervised Word Segmentation from Speech with Attention
We present a first attempt to perform attentional word segmentation directly from the speech signal, with the final goal to automatically identify lexical units in a low-resource, unwritten language (UL). Our methodology…
Acoustic Unit DiscoveryMachine TranslationSegmentationTranslationNew bilingual speech databases for audio diarization
This paper describes the process of collecting and recording two new bilingual speech databases in Spanish and Basque. They are designed primarily for speaker diarization in two different application domains: broadcast n…
speaker-diarizationSpeaker DiarizationSpeaker RecognitionThe IFCASL Corpus of French and German Non-native and Native Read Speech
The IFCASL corpus is a French-German bilingual phonetic learner corpus designed, recorded and annotated in a project on individualized feedback in computer-assisted spoken language learning. The motivation for setting up…
ArzEn: A Speech Corpus for Code-switched Egyptian Arabic-English
In this paper, we present our ArzEn corpus, an Egyptian Arabic-English code-switching (CS) spontaneous speech corpus. The corpus is collected through informal interviews with 38 Egyptian bilingual university students and…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Sentencespeech-recognition+1How Good is Automatic Segmentation as a Multimodal Discourse Annotation Aid?
Collaborative problem solving (CPS) in teams is tightly coupled with the creation of shared meaning between participants in a situated, collaborative task. In this work, we assess the quality of different utterance segme…
Segmentation