paper-with-me

Papers

Transport-Oriented Feature Aggregation for Speaker Embedding Learning

2022-06-26 · Yusheng Tian, Jingyu Li, Tan Lee

Pooling is needed to aggregate frame-level features into utterance-level representations for speaker modeling. Given the success of statistics-based pooling methods, we hypothesize that speaker characteristics are well represented in the statistical distribution over the pre-aggregation layer's output, and propose to use transport-oriented feature aggregation for deriving speaker embeddings. The aggregated representation encodes the geometric structure of the underlying feature distribution, which is expected to contain valuable speaker-specific information that may not be represented by the commonly used statistical measures like mean and variance. The original transport-oriented feature aggregation is also extended to a weighted-frame version to incorporate the attention mechanism. Experiments on speaker verification with the Voxceleb dataset show improvement over statistics pooling and its attentive variant.

📄 PDF Abstract BibTeX arXiv:2206.12857

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker Verification

Similar Papers 제목 키워드 기반

Improving Multi-Scale Aggregation Using Feature Pyramid Module for Robust Speaker Verification of Variable-Duration Utterances

2020-04-07 · Youngmoon Jung, Seong Min Kye, Yeunju Choi, Myunghun Jung 외

Currently, the most widely used approach for speaker verification is the deep speaker embedding learning. In this approach, we obtain a speaker embedding vector by pooling single-scale features that are extracted from th…

Speaker VerificationText-Independent Speaker Verification

Self-Attentive Multi-Layer Aggregation with Feature Recalibration and Normalization for End-to-End Speaker Verification System

2020-07-27 · Soonshin Seo, Ji-Hwan Kim

One of the most important parts of an end-to-end speaker verification system is the speaker embedding generation. In our previous paper, we reported that shortcut connections-based multi-layer aggregation improves the re…

Speaker Verification

X-Vectors with Multi-Scale Aggregation for Speaker Diarization

2021-05-16 · Myungjong Kim, Vijendra Raj Apsingekar, Divya Neelagiri

Speaker diarization is the process of labeling different speakers in a speech signal. Deep speaker embeddings are generally extracted from short speech segments and clustered to determine the segments belong to same spea…

speaker-diarizationSpeaker DiarizationSpeaker Verification

MCSAE: Masked Cross Self-Attentive Encoding for Speaker Embedding

2020-01-28 · Soonshin Seo, Ji-Hwan Kim

In general, a self-attention mechanism has been applied for speaker embedding encoding. Previous studies focused on training the self-attention in a high-level layer, such as the last pooling layer. However, the effect o…

Speaker Verification

High-resolution embedding extractor for speaker diarisation

2022-11-08 · Hee-Soo Heo, Youngki Kwon, Bong-Jin Lee, You Jin Kim 외

Speaker embedding extractors significantly influence the performance of clustering-based speaker diarisation systems. Conventionally, only one embedding is extracted from each speech segment. However, because of the slid…

Vocal Bursts Intensity Prediction