paper-with-me

홈 › Papers

Phonetic-attention scoring for deep speaker features in speaker verification

2018-11-08 · Lantian Li, Zhiyuan Tang, Ying Shi, Dong Wang

Recent studies have shown that frame-level deep speaker features can be derived from a deep neural network with the training target set to discriminate speakers by a short speech segment. By pooling the frame-level features, utterance-level representations, called d-vectors, can be derived and used in the automatic speaker verification (ASV) task. This simple average pooling, however, is inherently sensitive to the phonetic content of the utterance. An interesting idea borrowed from machine translation is the attention-based mechanism, where the contribution of an input word to the translation at a particular time is weighted by an attention score. This score reflects the relevance of the input word and the present translation. We can use the same idea to align utterances with different phonetic contents. This paper proposes a phonetic-attention scoring approach for d-vector systems. By this approach, an attention score is computed for each frame pair. This score reflects the similarity of the two frames in phonetic content, and is used to weigh the contribution of this frame pair in the utterance-based scoring. This new scoring approach emphasizes the frame pairs with similar phonetic contents, which essentially provides a soft alignment for utterances with any phonetic contents. Experimental results show that compared with the naive average pooling, this phonetic-attention scoring approach can deliver consistent performance improvement in ASV tasks of both text-dependent and text-independent.

📄 PDF Abstract BibTeX arXiv:1811.03255

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationSpeaker VerificationTranslation

Similar Papers 제목 키워드 기반

End-to-End Attention based Text-Dependent Speaker Verification

2017-01-03 · Shi-Xiong Zhang, Zhuo Chen, Yong Zhao, Jinyu Li 외

A new type of End-to-End system for text-dependent speaker verification is presented in this paper. Previously, using the phonetically discriminative/speaker discriminative DNNs as feature extractors for speaker verifica…

Speaker VerificationText-Dependent Speaker Verification

PDAF: A Phonetic Debiasing Attention Framework For Speaker Verification

2024-09-09 · Massa Baali, Abdulhamid Aldoobi, Hira Dhamyal, Rita Singh 외

Speaker verification systems are crucial for authenticating identity through voice. Traditionally, these systems focus on comparing feature vectors, overlooking the speech's content. However, this paper challenges this b…

Speaker Verification

Discrete Unit based Masking for Improving Disentanglement in Voice Conversion

2024-09-17 · Philip H. Lee, Ismail Rasim Ulgen, Berrak Sisman

Voice conversion (VC) aims to modify the speaker's identity while preserving the linguistic content. Commonly, VC methods use an encoder-decoder architecture, where disentangling the speaker's identity from linguistic in…

DecoderDisentanglementVoice Conversion

ExPO: Explainable Phonetic Trait-Oriented Network for Speaker Verification

2025-01-10 · Yi Ma, Shuai Wang, Tianchi Liu, Haizhou Li

In speaker verification, we use computational method to verify if an utterance matches the identity of an enrolled speaker. This task is similar to the manual task of forensic voice comparison, where linguistic analysis …

Speaker Verification

Phonetic-aware speaker embedding for far-field speaker verification

2023-11-27 · Zezhong Jin, Youzhi Tu, Man-Wai Mak

When a speaker verification (SV) system operates far from the sound sourced, significant challenges arise due to the interference of noise and reverberation. Studies have shown that incorporating phonetic information int…

Speaker RecognitionSpeaker Verificationspeech-recognitionSpeech Recognition