paper-with-me

홈 › Papers

Speech Decomposition Based on a Hybrid Speech Model and Optimal Segmentation

2021-05-04 · Alfredo Esquivel Jaramillo, Jesper Kjær Nielsen, Mads Græsbøll Christensen

In a hybrid speech model, both voiced and unvoiced components can coexist in a segment. Often, the voiced speech is regarded as the deterministic component, and the unvoiced speech and additive noise are the stochastic components. Typically, the speech signal is considered stationary within fixed segments of 20-40 ms, but the degree of stationarity varies over time. For decomposing noisy speech into its voiced and unvoiced components, a fixed segmentation may be too crude, and we here propose to adapt the segment length according to the signal local characteristics. The segmentation relies on parameter estimates of a hybrid speech model and the maximum a posteriori (MAP) and log-likelihood criteria as rules for model selection among the possible segment lengths, for voiced and unvoiced speech, respectively. Given the optimal segmentation markers and the estimated statistics, both components are estimated using linear filtering. A codebook-based approach differentiates between unvoiced speech and noise. A better extraction of the components is possible by taking into account the adaptive segmentation, compared to a fixed one. Also, a lower distortion for voiced speech and higher segSNR for both components is possible, as compared to other decomposition methods.

📄 PDF Abstract BibTeX arXiv:2105.01302

Code (0)

등록된 구현이 없습니다.

Tasks

Model SelectionSegmentation

Similar Papers 제목 키워드 기반

Beyond Voice Activity Detection: Hybrid Audio Segmentation for Direct Speech Translation

2021-04-23 · ICNLSP 2021 11 · Marco Gaido, Matteo Negri, Mauro Cettolo, Marco Turchi

The audio segmentation mismatch between training data and those seen at run-time is a major problem in direct speech translation. Indeed, while systems are usually trained on manually segmented corpora, in real use cases…

Action DetectionActivity DetectionSegmentationTranslation

Speech Segmentation Optimization using Segmented Bilingual Speech Corpus for End-to-end Speech Translation

2022-03-29 · Ryo Fukuda, Katsuhito Sudoh, Satoshi Nakamura

Speech segmentation, which splits long speech into short segments, is essential for speech translation (ST). Popular VAD tools like WebRTC VAD have generally relied on pause-based segmentation. Unfortunately, pauses in s…

Binary ClassificationSegmentationSentenceTranslation

SHAS: Approaching optimal Segmentation for End-to-End Speech Translation

2022-02-09 · Ioannis Tsiamas, Gerard I. Gállego, José A. R. Fonollosa, Marta R. Costa-jussà

Speech translation models are unable to directly process long audios, like TED talks, which have to be split into shorter segments. Speech translation datasets provide manual segmentations of the audios, which are not av…

SegmentationSpeech-to-Text TranslationTranslation

Speech segmentation using multilevel hybrid filters

2022-02-24 · Marcos Faundez-Zanuy, Francesc Vallverdu-Bayes

A novel approach for speech segmentation is proposed, based on Multilevel Hybrid (mean/min) Filters (MHF) with the following features: An accurate transition location. Good performance in noisy environments (gaussian and…

Segmentation

WaDeNet: Wavelet Decomposition based CNN for Speech Processing

2020-11-11 · Prithvi Suresh, Abhijith Ragav

Existing speech processing systems consist of different modules, individually optimized for a specific task such as acoustic modelling or feature extraction. In addition to not assuring optimality of the system, the disj…

Acoustic ModellingEmotion Recognition