paper-with-me

홈 › Papers

Deep Representation Learning in Speech Processing: Challenges, Recent Advances, and Future Trends

2020-01-02 · Siddique Latif, Rajib Rana, Sara Khalifa, Raja Jurdak, Junaid Qadir, Björn W. Schuller

Research on speech processing has traditionally considered the task of designing hand-engineered acoustic features (feature engineering) as a separate distinct problem from the task of designing efficient machine learning (ML) models to make prediction and classification decisions. There are two main drawbacks to this approach: firstly, the feature engineering being manual is cumbersome and requires human knowledge; and secondly, the designed features might not be best for the objective at hand. This has motivated the adoption of a recent trend in speech community towards utilisation of representation learning techniques, which can learn an intermediate representation of the input signal automatically that better suits the task at hand and hence lead to improved performance. The significance of representation learning has increased with advances in deep learning (DL), where the representations are more useful and less dependent on human knowledge, making it very conducive for tasks like classification, prediction, etc. The main contribution of this paper is to present an up-to-date and comprehensive survey on different techniques of speech representation learning by bringing together the scattered research across three distinct research areas including Automatic Speech Recognition (ASR), Speaker Recognition (SR), and Speaker Emotion Recognition (SER). Recent reviews in speech have been conducted for ASR, SR, and SER, however, none of these has focused on the representation learning from speech -- a gap that our survey aims to bridge.

📄 PDF Abstract BibTeX arXiv:2001.00378

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion RecognitionFeature EngineeringRepresentation LearningSpeaker Recognitionspeech-recognitionSpeech RecognitionSpeech Representation LearningSurvey

Similar Papers 제목 키워드 기반

Towards a Common Speech Analysis Engine

2022-03-01 · Hagai Aronowitz, Itai Gat, Edmilson Morais, Weizhong Zhu 외

Recent innovations in self-supervised representation learning have led to remarkable advances in natural language processing. That said, in the speech processing domain, self-supervised representation learning-based syst…

Emotion RecognitionLanguage IdentificationRepresentation Learning

Recent Advances in End-to-End Simultaneous Speech Translation

2024-06-01 · Xiaoqian Liu, Guoqiang Hu, Yangfan Du, Erfeng He 외

Simultaneous speech translation (SimulST) is a demanding task that involves generating translations in real-time while continuously processing speech input. This paper offers a comprehensive overview of the recent develo…

Translation

Speech Tokenizer is Key to Consistent Representation

2025-07-09 · Wonjin Jung, Sungil Kang, Dong-Yeon Cho arxiv

Speech tokenization is crucial in digital speech processing, converting continuous speech signals into discrete units for various computational tasks. This paper introduces a novel speech tokenizer with broad applicabili…

Emotion RecognitionVoice Conversion

A Review of Deep Learning Techniques for Speech Processing

2023-04-30 · Ambuj Mehrish, Navonil Majumder, Rishabh Bhardwaj, Rada Mihalcea 외

The field of speech processing has undergone a transformative shift with the advent of deep learning. The use of multiple processing layers has enabled the creation of models capable of extracting intricate features from…

Automatic Speech RecognitionDeep LearningEmotion Recognitionspeech-recognition+5

Temporal-Spatial Neural Filter: Direction Informed End-to-End Multi-channel Target Speech Separation

2020-01-02 · Rongzhi Gu, Yuexian Zou

Target speech separation refers to extracting the target speaker's speech from mixed signals. Despite the recent advances in deep learning based close-talk speech separation, the applications to real-world are still an o…

Speech Separation