paper-with-me

홈 › Papers

Speaker Embeddings as Individuality Proxy for Voice Stress Detection

2023-06-09 · Zihan Wu, Neil Scheidwasser-Clow, Karl El Hajal, Milos Cernak

Since the mental states of the speaker modulate speech, stress introduced by cognitive or physical loads could be detected in the voice. The existing voice stress detection benchmark has shown that the audio embeddings extracted from the Hybrid BYOL-S self-supervised model perform well. However, the benchmark only evaluates performance separately on each dataset, but does not evaluate performance across the different types of stress and different languages. Moreover, previous studies found strong individual differences in stress susceptibility. This paper presents the design and development of voice stress detection, trained on more than 100 speakers from 9 language groups and five different types of stress. We address individual variabilities in voice stress analysis by adding speaker embeddings to the hybrid BYOL-S features. The proposed method significantly improves voice stress detection performance with an input audio length of only 3-5 seconds.

📄 PDF Abstract BibTeX arXiv:2306.05915

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Creating Personalized Synthetic Voices from Post-Glossectomy Speech with Guided Diffusion Models

2023-05-27 · Yusheng Tian, Guangyan Zhang, Tan Lee

This paper is about developing personalized speech synthesis systems with recordings of mildly impaired speech. In particular, we consider consonant and vowel alterations resulted from partial glossectomy, the surgical r…

Speech SynthesisVoice Conversion

V2S attack: building DNN-based voice conversion from automatic speaker verification

2019-08-05 · Taiki Nakamura, Yuki Saito, Shinnosuke Takamichi, Yusuke Ijima 외

This paper presents a new voice impersonation attack using voice conversion (VC). Enrolling personal voices for automatic speaker verification (ASV) offers natural and flexible biometric authentication systems. Basically…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speaker Verificationspeech-recognition+2

An objective evaluation of the effects of recording conditions and speaker characteristics in multi-speaker deep neural speech synthesis

2021-06-03 · Beata Lorincz, Adriana Stan, Mircea Giurgiu

Multi-speaker spoken datasets enable the creation of text-to-speech synthesis (TTS) systems which can output several voice identities. The multi-speaker (MSPK) scenario also enables the use of fewer training samples per …

Speaker VerificationSpeech Synthesistext-to-speechText to Speech+1

Creating Personalized Synthetic Voices from Articulation Impaired Speech Using Augmented Reconstruction Loss

2024-01-08 · Yusheng Tian, Jingyu Li, Tan Lee

This research is about the creation of personalized synthetic voices for head and neck cancer survivors. It is focused particularly on tongue cancer patients whose speech might exhibit severe articulation impairment. Our…

ZSDEVC: Zero-Shot Diffusion-based Emotional Voice Conversion with Disentangled Mechanism

2024-09-05 · Hsing-Hang Chou, Yun-Shao Lin, Ching-Chin Sung, Yu Tsao 외

The human voice conveys not just words but also emotional states and individuality. Emotional voice conversion (EVC) modifies emotional expressions while preserving linguistic content and speaker identity, improving appl…

Emotion ClassificationVoice Conversion