paper-with-me

Papers

Binary Speaker Embedding

2015-10-20 · Lantian Li, Dong Wang, Chao Xing, Kaimin Yu, Thomas Fang Zheng

The popular i-vector model represents speakers as low-dimensional continuous vectors (i-vectors), and hence it is a way of continuous speaker embedding. In this paper, we investigate binary speaker embedding, which transforms i-vectors to binary vectors (codes) by a hash function. We start from locality sensitive hashing (LSH), a simple binarization approach where binary codes are derived from a set of random hash functions. A potential problem of LSH is that the randomly sampled hash functions might be suboptimal. We therefore propose an improved Hamming distance learning approach, where the hash function is learned by a variable-sized block training that projects each dimension of the original i-vectors to variable-sized binary codes independently. Our experiments show that binary speaker embedding can deliver competitive or even better results on both speaker verification and identification tasks, while the memory usage and the computation cost are significantly reduced.

📄 PDF Abstract BibTeX arXiv:1510.05937

Code (0)

등록된 구현이 없습니다.

Tasks

BinarizationSpeaker Verification

Similar Papers 제목 키워드 기반

Ordered and Binary Speaker Embedding

2023-05-25 · Jiaying Wang, Xianglong Wang, Namin Wang, Lantian Li 외

Modern speaker recognition systems represent utterances by embedding vectors. Conventional embedding vectors are dense and non-structural. In this paper, we propose an ordered binary embedding approach that sorts the dim…

ClusteringRetrievalSpeaker IdentificationSpeaker Recognition

Speaker Embedding-aware Neural Diarization: an Efficient Framework for Overlapping Speech Diarization in Meeting Scenarios

2022-03-18 · Zhihao Du, Shiliang Zhang, Siqi Zheng, Zhijie Yan

Overlapping speech diarization has been traditionally treated as a multi-label classification problem. In this paper, we reformulate this task as a single-label prediction problem by encoding multiple binary labels into …

Action DetectionActivity DetectionMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATION+1

A Novel Automatic Framework for Speaker Drift Detection in Synthesized Speech

2026-04-07 · Jia-Hong Huang, Seulgi Kim, Yi Chieh Liu, Yixian Shen 외 arxiv

Recent diffusion-based text-to-speech (TTS) models achieve high naturalness and expressiveness, yet often suffer from speaker drift, a subtle, gradual shift in perceived speaker identity within a single utterance. This u…

Binary Classification

Combining Automatic Speaker Verification and Prosody Analysis for Synthetic Speech Detection

2022-10-31 · Luigi Attorresi, Davide Salvi, Clara Borrelli, Paolo Bestagini 외

The rapid spread of media content synthesis technology and the potentially damaging impact of audio and video deepfakes on people's lives have raised the need to implement systems able to detect these forgeries automatic…

Audio CompressionFace SwappingRhythmSpeaker Verification+4

The UPC Speaker Verification System Submitted to VoxCeleb Speaker Recognition Challenge 2020 (VoxSRC-20)

2020-10-27

This report describes the submission from Technical University of Catalonia (UPC) to the VoxCeleb Speaker Recognition Challenge (VoxSRC-20) at Interspeech 2020. The final submission is a combination of three systems. Sys…

Binary ClassificationSpeaker RecognitionSpeaker VerificationTriplet