paper-with-me

Papers

STC Speaker Recognition Systems for the VOiCES From a Distance Challenge

2019-04-12 · Sergey Novoselov, Aleksei Gusev, Artem Ivanov, Timur Pekhovsky, Andrey Shulipa, Galina Lavrentyeva, Vladimir Volokhov, Alexandr Kozlov

This paper presents the Speech Technology Center (STC) speaker recognition (SR) systems submitted to the VOiCES From a Distance challenge 2019. The challenge's SR task is focused on the problem of speaker recognition in single channel distant/far-field audio under noisy conditions. In this work we investigate different deep neural networks architectures for speaker embedding extraction to solve the task. We show that deep networks with residual frame level connections outperform more shallow architectures. Simple energy based speech activity detector (SAD) and automatic speech recognition (ASR) based SAD are investigated in this work. We also address the problem of data preparation for robust embedding extractors training. The reverberation for the data augmentation was performed using automatic room impulse response generator. In our systems we used discriminatively trained cosine similarity metric learning model as embedding backend. Scores normalization procedure was applied for each individual subsystem we used. Our final submitted systems were based on the fusion of different subsystems. The results obtained on the VOiCES development and evaluation sets demonstrate effectiveness and robustness of the proposed systems when dealing with distant/far-field audio under noisy conditions.

📄 PDF Abstract BibTeX arXiv:1904.06093

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationMetric LearningRoom Impulse Response (RIR)Speaker Recognitionspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Who is Authentic Speaker

2024-04-30 · Qiang Huang

Voice conversion (VC) using deep learning technologies can now generate high quality one-to-many voices and thus has been used in some practical application fields, such as entertainment and healthcare. However, voice co…

Speaker RecognitionVoice Conversion

Individualized Conditioning and Negative Distances for Speaker Separation

2022-10-12 · Tao Sun, Nidal Abuhajar, Shuyu Gong, Zhewei Wang 외

Speaker separation aims to extract multiple voices from a mixed signal. In this paper, we propose two speaker-aware designs to improve the existing speaker separation solutions. The first model is a speaker conditioning …

Speaker SeparationTriplet

BUT VOiCES 2019 System Description

2019-07-13 · Hossein Zeinali, Pavel Matějka, Ladislav Mošner, Oldřich Plchot 외

This is a description of our effort in VOiCES 2019 Speaker Recognition challenge. All systems in the fixed condition are based on the x-vector paradigm with different features and DNN topologies. The single best system r…

Speaker Recognition

SLMIA-SR: Speaker-Level Membership Inference Attacks against Speaker Recognition Systems

2023-09-14 · Guangke Chen, Yedi Zhang, Fu Song

Membership inference attacks allow adversaries to determine whether a particular example was contained in the model's training dataset. While previous works have confirmed the feasibility of such attacks in various appli…

Feature EngineeringInference AttackMembership Inference AttackSpeaker Recognition

AS2T: Arbitrary Source-To-Target Adversarial Attack on Speaker Recognition Systems

2022-06-07 · Guangke Chen, Zhe Zhao, Fu Song, Sen Chen 외

Recent work has illuminated the vulnerability of speaker recognition systems (SRSs) against adversarial attacks, raising significant security concerns in deploying SRSs. However, they considered only a few settings (e.g.…

Adversarial AttackSpeaker Recognition