openFEAT: Improving Speaker Identification by Open-set Few-shot Embedding Adaptation with Transformer
Household speaker identification with few enrollment utterances is an important yet challenging problem, especially when household members share similar voice characteristics and room acoustics. A common embedding space learned from a large number of speakers is not universally applicable for the optimal identification of every speaker in a household. In this work, we first formulate household speaker identification as a few-shot open-set recognition task and then propose a novel embedding adaptation framework to adapt speaker representations from the given universal embedding space to a household-specific embedding space using a set-to-set function, yielding better household speaker identification performance. With our algorithm, Open-set Few-shot Embedding Adaptation with Transformer (openFEAT), we observe that the speaker identification equal error rate (IEER) on simulated households with 2 to 7 hard-to-discriminate speakers is reduced by 23% to 31% relative.
Code (0)
등록된 구현이 없습니다.
Tasks
Open Set LearningSpeaker IdentificationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Centroid-based deep metric learning for speaker recognition
Speaker embedding models that utilize neural networks to map utterances to a space where distances reflect similarity between speakers have driven recent progress in the speaker recognition task. However, there is still …
Few-Shot Image ClassificationFew-Shot LearningGeneral Classificationimage-classification+5Improved Relation Networks for End-to-End Speaker Verification and Identification
Speaker identification systems in a real-world scenario are tasked to identify a speaker amongst a set of enrolled speakers given just a few samples for each enrolled speaker. This paper demonstrates the effectiveness of…
Meta-LearningRelationSpeaker IdentificationSpeaker VerificationEnhancing Open-Set Speaker Identification through Rapid Tuning with Speaker Reciprocal Points and Negative Sample
This paper introduces a novel framework for open-set speaker identification in household environments, playing a crucial role in facilitating seamless human-computer interactions. Addressing the limitations of current sp…
Speaker IdentificationSpeaker RecognitionAccentBox: Towards High-Fidelity Zero-Shot Accent Generation
While recent Zero-Shot Text-to-Speech (ZS-TTS) models have achieved high naturalness and speaker similarity, they fall short in accent fidelity and control. To address this issue, we propose zero-shot accent generation t…
text-to-speechText to SpeechSymmetric Saliency-based Adversarial Attack To Speaker Identification
Adversarial attack approaches to speaker identification either need high computational cost or are not very effective, to our knowledge. To address this issue, in this paper, we propose a novel generation-network-based a…
Adversarial AttackDecoderSpeaker Identification