paper-with-me

홈 › Papers

Speaker Identification from emotional and noisy speech data using learned voice segregation and Speech VGG

2022-10-23 · Shibani Hamsa, Ismail Shahin, Youssef Iraqi, Ernesto Damiani, Naoufel Werghi

Speech signals are subjected to more acoustic interference and emotional factors than other signals. Noisy emotion-riddled speech data is a challenge for real-time speech processing applications. It is essential to find an effective way to segregate the dominant signal from other external influences. An ideal system should have the capacity to accurately recognize required auditory events from a complex scene taken in an unfavorable situation. This paper proposes a novel approach to speaker identification in unfavorable conditions such as emotion and interference using a pre-trained Deep Neural Network mask and speech VGG. The proposed model obtained superior performance over the recent literature in English and Arabic emotional speech data and reported an average speaker identification rate of 85.2\%, 87.0\%, and 86.6\% using the Ryerson audio-visual dataset (RAVDESS), speech under simulated and actual stress (SUSAS) dataset and Emirati-accented Speech dataset (ESD) respectively.

📄 PDF Abstract BibTeX arXiv:2210.12701

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker Identification

Methods 이 논문이 사용한 방법론

Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…

Similar Papers 제목 키워드 기반

CASA-Based Speaker Identification Using Cascaded GMM-CNN Classifier in Noisy and Emotional Talking Conditions

2021-02-11 · Ali Bou Nassif, Ismail Shahin, Shibani Hamsa, Nawel Nemmour 외

This work aims at intensifying text-independent speaker identification performance in real application situations such as noisy and emotional talking conditions. This is achieved by incorporating two different modules: a…

Emotion RecognitionSpeaker Identification

End-to-end Recurrent Denoising Autoencoder Embeddings for Speaker Identification

2020-03-13 · Esther Rituerto-González, Carmen Peláez-Moreno

Speech 'in-the-wild' is a handicap for speaker recognition systems due to the variability induced by real-life conditions, such as environmental noise and the emotional state of the speaker. Taking advantage of the princ…

Data AugmentationDenoisingRepresentation LearningSpeaker Identification+1

Three-Stage Speaker Verification Architecture in Emotional Talking Environments

2018-09-03 · Ismail Shahin, Ali Bou Nassif

Speaker verification performance in neutral talking environment is usually high, while it is sharply decreased in emotional talking environments. This performance degradation in emotional environments is due to the probl…

Speaker Verification

Exploring VQ-VAE with Prosody Parameters for Speaker Anonymization

2024-09-24 · Sotheara Leang, Anderson Augusma, Eric Castelli, Frédérique Letué 외

Human speech conveys prosody, linguistic content, and speaker identity. This article investigates a novel speaker anonymization approach using an end-to-end network based on a Vector-Quantized Variational Auto-Encoder (V…

DecoderSpeaker anonymizationSpeaker Identification

Characteristic-Specific Partial Fine-Tuning for Efficient Emotion and Speaker Adaptation in Codec Language Text-to-Speech Models

2025-01-24 · Tianrui Wang, Meng Ge, Cheng Gong, Chunyu Qiang 외

Recently, emotional speech generation and speaker cloning have garnered significant interest in text-to-speech (TTS). With the open-sourcing of codec language TTS models trained on massive datasets with large-scale param…

Emotion ClassificationSpeaker Identificationtext-to-speechText to Speech