paper-with-me

홈 › Papers

A Subband-Based SVM Front-End for Robust ASR

2013-12-24 · Jibran Yousafzai, Zoran Cvetkovic, Peter Sollich, Matthew Ager

This work proposes a novel support vector machine (SVM) based robust automatic speech recognition (ASR) front-end that operates on an ensemble of the subband components of high-dimensional acoustic waveforms. The key issues of selecting the appropriate SVM kernels for classification in frequency subbands and the combination of individual subband classifiers using ensemble methods are addressed. The proposed front-end is compared with state-of-the-art ASR front-ends in terms of robustness to additive noise and linear filtering. Experiments performed on the TIMIT phoneme classification task demonstrate the benefits of the proposed subband based SVM front-end: it outperforms the standard cepstral front-end in the presence of noise and linear filtering for signal-to-noise ratio (SNR) below 12-dB. A combination of the proposed front-end with a conventional front-end such as MFCC yields further improvements over the individual front ends across the full range of noise levels.

📄 PDF Abstract BibTeX arXiv:1401.3322

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)General Classificationspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

SVM A Support Vector Machine, or SVM, is a non-parametric supervised learning model. For non-linear classification and regression, they utilise the kernel trick to map inputs…

Similar Papers 제목 키워드 기반

Digital Prototype Filter Alternatives for Processing Frequency-Stacked Mobile Subbands Deploying a Single ADC for Beamforming Satellites

2024-01-05 · Adem Coskun, Sevket Cetinsel, Izzet Kale, Robert Hughes 외

This article presents a two-stage approach for the processing of frequency-stacked mobile subbands. The frequency stacking is performed in the analog domain to enable the use of a wideband analog-to-digital converter (AD…

RatioWaveNet: A Learnable RDWT Front-End for Robust and Interpretable EEG Motor-Imagery Classification

2025-10-22 · Marco Siino, Giuseppe Bonomo, Rosario Sorbello, Ilenia Tinnirello arxiv

Brain-computer interfaces (BCIs) based on motor imagery (MI) translate covert movement intentions into actionable commands, yet reliable decoding from non-invasive EEG remains challenging due to nonstationarity, low SNR,…

Joint Phase Time Array: Opportunities, Challenges and System Design Considerations

2024-12-02 · Young-Han Nam, Ahmad AlAmmouri, Jianhua Mo, Jianzhong Chalrie Zhang

This paper presents a novel approach to designing millimeter-wave (mmWave) cellular communication systems, based on joint phase time array (JPTA) radio frequency (RF) frontend architecture. JPTA architecture comprises ti…

A Fully Time-domain Neural Model for Subband-based Speech Synthesizer

2018-10-12 · Azam Rabiee, Soo-Young Lee

This paper introduces a deep neural network model for subband-based speech synthesizer. The model benefits from the short bandwidth of the subband signals to reduce the complexity of the time-domain speech generator. We …

text-to-speechText to Speech

A Fully Time-domain Neural Model for Subband-based Speech Synthesizer

2018-10-22 · Anonymous

—This paper introduces a deep neural network model for subband-based speech synthesizer. The model benefits from the short bandwidth of the subband signals to reduce the complexity of the time-do…

text-to-speechText to Speech