paper-with-me

Papers

A Novel Speech Feature Fusion Algorithm for Text-Independent Speaker Recognition

2022-12-01 · Biao Ma, Chengben Xu, Ye Zhang

A novel speech feature fusion algorithm with independent vector analysis (IVA) and parallel convolutional neural network (PCNN) is proposed for text-independent speaker recognition. Firstly, some different feature types, such as the time domain (TD) features and the frequency domain (FD) features, can be extracted from a speaker's speech, and the TD and the FD features can be considered as the linear mixtures of independent feature components (IFCs) with an unknown mixing system. To estimate the IFCs, the TD and the FD features of the speaker's speech are concatenated to build the TD and the FD feature matrix, respectively. Then, a feature tensor of the speaker's speech is obtained by paralleling the TD and the FD feature matrix. To enhance the dependence on different feature types and remove the redundancies of the same feature type, the independent vector analysis (IVA) can be used to estimate the IFC matrices of TD and FD features with the feature tensor. The IFC matrices are utilized as the input of the PCNN to extract the deep features of the TD and FD features, respectively. The deep features can be integrated to obtain the fusion feature of the speaker's speech. Finally, the fusion feature of the speaker's speech is employed as the input of a deep convolutional neural network (DCNN) classifier for speaker recognition. The experimental results show the effectiveness and performances of the proposed speaker recognition system.

📄 PDF Abstract BibTeX arXiv:2212.00329

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker RecognitionText-Independent Speaker Recognition

Similar Papers 제목 키워드 기반

Independent Feature Enhanced Crossmodal Fusion for Match-Mismatch Classification of Speech Stimulus and EEG Response

2024-10-19 · Shitong Fan, Wenbo Wang, Feiyang Xiao, Shiheng Zhang 외

It is crucial for auditory attention decoding to classify matched and mismatched speech stimuli with corresponding EEG responses by exploring their relationship. However, existing methods often adopt two independent netw…

EEGEeg Decoding

Cross-modal Audio-visual Co-learning for Text-independent Speaker Verification

2023-02-22 · Meng Liu, Kong Aik Lee, Longbiao Wang, Hanyi Zhang 외

Visual speech (i.e., lip motion) is highly related to auditory speech due to the co-occurrence and synchronization in speech production. This paper investigates this correlation and proposes a cross-modal speech co-learn…

Speaker VerificationText-Independent Speaker Verification

VoiceExtender: Short-utterance Text-independent Speaker Verification with Guided Diffusion Model

2023-10-07 · Yayun He, Zuheng Kang, Jianzong Wang, Junqing Peng 외

Speaker verification (SV) performance deteriorates as utterances become shorter. To this end, we propose a new architecture called VoiceExtender which provides a promising solution for improving SV performance when handl…

Speaker VerificationText-Independent Speaker Verification

Multimodal Emotion Recognition with Transformer-Based Self Supervised Feature Fusion

2020-10-27 · Shamane Siriwardhana ; Tharindu Kaluarachchi ; Mark Billinghurst ; Suranga Nanayakkara

Emotion Recognition is a challenging research area given its complex nature, and humans express emotional cues across various modalities such as language, facial expressions, and speech. Representation and fusion of feat…

Emotion RecognitionMultimodal Deep LearningMultimodal Emotion RecognitionMultimodal Sentiment Analysis+2

Speaker Independent Continuous Speech to Text Converter for Mobile Application

2013-07-19 · R. Sandanalakshmi, P. Abinaya Viji, M. Kiruthiga, M. Manjari 외

An efficient speech to text converter for mobile application is presented in this work. The prime motive is to formulate a system which would give optimum performance in terms of complexity, accuracy, delay and memory re…

Action DetectionActivity DetectionSpeech-to-Text