paper-with-me

Papers

Speech Segmentation Optimization using Segmented Bilingual Speech Corpus for End-to-end Speech Translation

2022-03-29 · Ryo Fukuda, Katsuhito Sudoh, Satoshi Nakamura

Speech segmentation, which splits long speech into short segments, is essential for speech translation (ST). Popular VAD tools like WebRTC VAD have generally relied on pause-based segmentation. Unfortunately, pauses in speech do not necessarily match sentence boundaries, and sentences can be connected by a very short pause that is difficult to detect by VAD. In this study, we propose a speech segmentation method using a binary classification model trained using a segmented bilingual speech corpus. We also propose a hybrid method that combines VAD and the above speech segmentation method. Experimental results revealed that the proposed method is more suitable for cascade and end-to-end ST systems than conventional segmentation methods. The hybrid approach further improved the translation performance.

📄 PDF Abstract BibTeX arXiv:2203.15479

Code (1)

wiseman/py-webrtcvad 공식 구현

Tasks

Binary ClassificationSegmentationSentenceTranslation

Similar Papers 제목 키워드 기반

Unsupervised Word Segmentation from Speech with Attention

2018-06-18 · Pierre Godard, Marcely Zanon-Boito, Lucas Ondel, Alexandre Berard 외

We present a first attempt to perform attentional word segmentation directly from the speech signal, with the final goal to automatically identify lexical units in a low-resource, unwritten language (UL). Our methodology…

Acoustic Unit DiscoveryMachine TranslationSegmentationTranslation

New bilingual speech databases for audio diarization

2014-05-01 · LREC 2014 5 · David Tavarez, Eva Navas, Daniel Erro, Ibon Saratxaga 외

This paper describes the process of collecting and recording two new bilingual speech databases in Spanish and Basque. They are designed primarily for speaker diarization in two different application domains: broadcast n…

speaker-diarizationSpeaker DiarizationSpeaker Recognition

The IFCASL Corpus of French and German Non-native and Native Read Speech

2016-05-01 · LREC 2016 5 · Juergen Trouvain, Anne Bonneau, Vincent Colotte, Camille Fauth 외

The IFCASL corpus is a French-German bilingual phonetic learner corpus designed, recorded and annotated in a project on individualized feedback in computer-assisted spoken language learning. The motivation for setting up…

ArzEn: A Speech Corpus for Code-switched Egyptian Arabic-English

2020-05-01 · LREC 2020 5 · Injy Hamed, Ngoc Thang Vu, Slim Abdennadher

In this paper, we present our ArzEn corpus, an Egyptian Arabic-English code-switching (CS) spontaneous speech corpus. The corpus is collected through informal interviews with 38 Egyptian bilingual university students and…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Sentencespeech-recognition+1

How Good is Automatic Segmentation as a Multimodal Discourse Annotation Aid?

2023-05-27 · Corbin Terpstra, Ibrahim Khebour, Mariah Bradford, Brett Wisniewski 외

Collaborative problem solving (CPS) in teams is tightly coupled with the creation of shared meaning between participants in a situated, collaborative task. In this work, we assess the quality of different utterance segme…

Segmentation