paper-with-me

Papers

Physiological-Physical Feature Fusion for Automatic Voice Spoofing Detection

2021-09-01 · Junxiao Xue, Hao Zhou, Yabo Wang

Speaker verification systems have been used in many production scenarios in recent years. Unfortunately, they are still highly prone to different kinds of spoofing attacks such as voice conversion and speech synthesis, etc. In this paper, we propose a new method base on physiological-physical feature fusion to deal with voice spoofing attacks. This method involves feature extraction, a densely connected convolutional neural network with squeeze and excitation block (SE-DenseNet), multi-scale residual neural network with squeeze and excitation block (SE-Res2Net) and feature fusion strategies. We first pre-trained a convolutional neural network using the speaker's voice and face in the video as surveillance signals. It can extract physiological features from speech. Then we use SE-DenseNet and SE-Res2Net to extract physical features. Such a densely connection pattern has high parameter efficiency and squeeze and excitation block can enhance the transmission of the feature. Finally, we integrate the two features into the SE-Densenet to identify the spoofing attacks. Experimental results on the ASVspoof 2019 data set show that our model is effective for voice spoofing detection. In the logical access scenario, our model improves the tandem decision cost function (t-DCF) and equal error rate (EER) scores by 4% and 7%, respectively, compared with other methods. In the physical access scenario, our model improved t-DCF and EER scores by 8% and 10%, respectively.

📄 PDF Abstract BibTeX arXiv:2109.00913

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker VerificationSpeech SynthesisVoice Conversion

Similar Papers 제목 키워드 기반

Hybrid Handcrafted and Learnable Audio Representation for Analysis of Speech Under Cognitive and Physical Load

2022-03-30 · Gasser Elbanna, Alice Biryukov, Neil Scheidwasser-Clow, Lara Orlandic 외

As a neurophysiological response to threat or adverse conditions, stress can affect cognition, emotion and behaviour with potentially detrimental effects on health in the case of sustained exposure. Since the affective c…

Representation Learning

Using Voice and Biofeedback to Predict User Engagement during Product Feedback Interviews

2021-04-06 · Alessio Ferrari, Thaide Huichapa, Paola Spoletini, Nicole Novielli 외

Capturing users' engagement is crucial for gathering feedback about the features of a software product. In a market-driven context, current approaches to collect and analyze users' feedback are based on techniques levera…

Impact of multiple modalities on emotion recognition: investigation into 3d facial landmarks, action units, and physiological data

2020-05-17 · Diego Fabiano, Manikandan Jaishanker, Shaun Canavan

To fully understand the complexities of human emotion, the integration of multiple physical features from different modalities can be advantageous. Considering this, we present an analysis of 3D facial data, action units…

Emotion Recognition

PHemoNet: A Multimodal Network for Physiological Signals

2024-09-13 · Eleonora Lopez, Aurelio Uncini, Danilo Comminiello

Emotion recognition is essential across numerous fields, including medical applications and brain-computer interface (BCI). Emotional responses include behavioral reactions, such as tone of voice and body movement, and c…

Brain Computer InterfaceEEGElectroencephalogram (EEG)Emotion Recognition+2

Coding Speech through Vocal Tract Kinematics

2024-06-18 · Cheol Jun Cho, Peter Wu, Tejas S. Prabhune, Dhruv Agarwal 외

Vocal tract articulation is a natural, grounded control space of speech production. The spatiotemporal coordination of articulators combined with the vocal source shapes intelligible speech sounds to enable effective spo…

Voice Conversion