paper-with-me

Papers

Triplet loss based embeddings for forensic speaker identification in Spanish

2021-02-24 · Emmanuel Maqueda, Javier Alvarez-Jimenez, Carlos Mena, Ivan Meza

With the advent of digital technology, it is more common that committed crimes or legal disputes involve some form of speech recording where the identity of a speaker is questioned [1]. In face of this situation, the field of forensic speaker identification has been looking to shed light on the problem by quantifying how much a speech recording belongs to a particular person in relation to a population. In this work, we explore the use of speech embeddings obtained by training a CNN using the triplet loss. In particular, we focus on the Spanish language which has not been extensively studies. We propose extracting the embeddings from speech spectrograms samples, then explore several configurations of such spectrograms, and finally, quantify the embeddings quality. We also show some limitations of our data setting which is predominantly composed by male speakers. At the end, we propose two approaches to calculate the Likelihood Radio given out speech embeddings and we show that triplet loss is a good alternative to create speech embeddings for forensic speaker identification.

📄 PDF Abstract BibTeX arXiv:2102.12564

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker IdentificationTriplet

Methods 이 논문이 사용한 방법론

Triplet Loss The goal of Triplet loss, in the context of Siamese Networks, is to maximize the joint probability among all score-pairs i.e. the product of all probabilities. By using its…

Similar Papers 제목 키워드 기반

Deep Speaker: an End-to-End Neural Speaker Embedding System

2017-05-05 · Chao Li, Xiaokong Ma, Bing Jiang, Xiangang Li 외

We present Deep Speaker, a neural speaker embedding system that maps utterances to a hypersphere where speaker similarity is measured by cosine similarity. The embeddings generated by Deep Speaker can be used for many ta…

ClusteringSpeaker IdentificationSpeaker RecognitionTriplet

Whisper Speaker Identification: Leveraging Pre-Trained Multilingual Transformers for Robust Speaker Embeddings

2025-03-13 · Jakaria Islam Emon, Md Abu Salek, Kazi Tamanna Alam

Speaker identification in multilingual settings presents unique challenges, particularly when conventional models are predominantly trained on English data. In this paper, we propose WSI (Whisper Speaker Identification),…

Speaker Identificationspeech-recognitionSpeech Recognition

Triplet Based Embedding Distance and Similarity Learning for Text-independent Speaker Verification

2019-08-06 · Zongze Ren, Zhiyong Chen, Shugong Xu

Speaker embeddings become growing popular in the text-independent speaker verification task. In this paper, we propose two improvements during the training stage. The improvements are both based on triplet cause the trai…

Speaker RecognitionSpeaker VerificationText-Independent Speaker VerificationTriplet

TristouNet: Triplet Loss for Speaker Turn Embedding

2016-09-14 · Hervé Bredin

TristouNet is a neural network architecture based on Long Short-Term Memory recurrent networks, meant to project speech sequences into a fixed-dimensional euclidean space. Thanks to the triplet loss paradigm used for tra…

Change DetectionTriplet

Triplet Entropy Loss: Improving The Generalisation of Short Speech Language Identification Systems

2020-12-03 · Ruan van der Merwe

We present several methods to improve the generalisation of language identification (LID) systems to new speakers and to new domains. These methods involve Spectral augmentation, where spectrograms are masked in the freq…

Language IdentificationSpeech Language IdentificationSpoken language identificationSpoken Language Understanding+1