paper-with-me

홈 › Papers

Text-Independent Speaker Identification Using Audio Looping With Margin Based Loss Functions

2025-09-26 · Elliot Q C Garcia, Nicéias Silva Vilela, Kátia Pires Nascimento do Sacramento, Tiago A. E. Ferreira arxiv

Speaker identification has become a crucial component in various applications, including security systems, virtual assistants, and personalized user experiences. In this paper, we investigate the effectiveness of CosFace Loss and ArcFace Loss for text-independent speaker identification using a Convolutional Neural Network architecture based on the VGG16 model, modified to accommodate mel spectrogram inputs of variable sizes generated from the Voxceleb1 dataset. Our approach involves implementing both loss functions to analyze their effects on model accuracy and robustness, where the Softmax loss function was employed as a comparative baseline. Additionally, we examine how the sizes of mel spectrograms and their varying time lengths influence model performance. The experimental results demonstrate superior identification accuracy compared to traditional Softmax loss methods. Furthermore, we discuss the implications of these findings for future research.

📄 PDF Abstract BibTeX arXiv:2509.22838

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker Identification

Similar Papers 제목 키워드 기반

Adaptive blind audio source extraction supervised by dominant speaker identification using x-vectors

2019-10-25

We propose a novel algorithm for adaptive blind audio source extraction. The proposed method is based on independent vector analysis and utilizes the auxiliary function optimization to achieve high convergence speed. The…

Speaker Identification

Rhythm Features for Speaker Identification

2025-06-07 · Nick Mehlman, Thomas Thebaud, Dani Byrd, Shri Narayanan

While deep learning models have demonstrated robust performance in speaker recognition tasks, they primarily rely on low-level audio features learned empirically from spectrograms or raw waveforms. However, prior work ha…

Deep LearningRhythmSpeaker IdentificationSpeaker Recognition

A Multi Level Data Fusion Approach for Speaker Identification on Telephone Speech

2014-06-27 · Imen Trabelsi, Dorra Ben Ayed

Several speaker identification systems are giving good performance with clean speech but are affected by the degradations introduced by noisy audio conditions. To deal with this problem, we investigate the use of complem…

Speaker Identification

CASA-Based Speaker Identification Using Cascaded GMM-CNN Classifier in Noisy and Emotional Talking Conditions

2021-02-11 · Ali Bou Nassif, Ismail Shahin, Shibani Hamsa, Nawel Nemmour 외

This work aims at intensifying text-independent speaker identification performance in real application situations such as noisy and emotional talking conditions. This is achieved by incorporating two different modules: a…

Emotion RecognitionSpeaker Identification

Automatic Voice Identification after Speech Resynthesis using PPG

2024-08-05 · Thibault Gaudier, Marie Tahon, Anthony Larcher, Yannick Estève

Speech resynthesis is a generic task for which we want to synthesize audio with another audio as input, which finds applications for media monitors and journalists.Among different tasks addressed by speech resynthesis, v…

ResynthesisSpeaker VerificationVoice Conversion