paper-with-me

Papers

Spatial Pyramid Encoding with Convex Length Normalization for Text-Independent Speaker Verification

2019-06-19 · Youngmoon Jung, Younggwan Kim, Hyungjun Lim, Yeunju Choi, Hoirin Kim

In this paper, we propose a new pooling method called spatial pyramid encoding (SPE) to generate speaker embeddings for text-independent speaker verification. We first partition the output feature maps from a deep residual network (ResNet) into increasingly fine sub-regions and extract speaker embeddings from each sub-region through a learnable dictionary encoding layer. These embeddings are concatenated to obtain the final speaker representation. The SPE layer not only generates a fixed-dimensional speaker embedding for a variable-length speech segment, but also aggregates the information of feature distribution from multi-level temporal bins. Furthermore, we apply deep length normalization by augmenting the loss function with ring loss. By applying ring loss, the network gradually learns to normalize the speaker embeddings using model weights themselves while preserving convexity, leading to more robust speaker embeddings. Experiments on the VoxCeleb1 dataset show that the proposed system using the SPE layer and ring loss-based deep length normalization outperforms both i-vector and d-vector baselines.

📄 PDF Abstract BibTeX arXiv:1906.08333

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker VerificationText-Independent Speaker Verification

Similar Papers 제목 키워드 기반

Attentional Pyramid Pooling of Salient Visual Residuals for Place Recognition

2021-01-01 · ICCV 2021 10 · Guohao Peng, Jun Zhang, Heshan Li, Danwei Wang

The core of visual place recognition (VPR) lies in how to identify task-relevant visual cues and embed them into discriminative representations. Focusing on these two points, we propose a novel encoding strategy name…

Visual Place Recognition

Unsupervised Representation Learning with Laplacian Pyramid Auto-encoders

2018-01-16 · Qilu Zhao, Zongmin Li

Scale-space representation has been popular in computer vision community due to its theoretical foundation. The motivation for generating a scale-space representation of a given data set originates from the basic observa…

Representation Learning

Beyond Spatial Pyramid Matching: Space-time Extended Descriptor for Action Recognition

2015-10-15 · Zhenzhong Lan, Alexander G. Hauptmann

We address the problem of generating video features for action recognition. The spatial pyramid and its variants have been very popular feature models due to their success in balancing spatial location encoding and spati…

Action RecognitionDiversityTemporal Action Localization

Linear Spatial Pyramid Matching Using Non-convex and non-negative Sparse Coding for Image Classification

2015-04-27 · Chengqiang Bao, Liangtian He, Yi-Lun Wang

Recently sparse coding have been highly successful in image classification mainly due to its capability of incorporating the sparsity of image representation. In this paper, we propose an improved sparse coding model bas…

General Classificationimage-classificationImage Classification

Deep Spatial Pyramid: The Devil is Once Again in the Details

2015-04-21 · Bin-Bin Gao, Xiu-Shen Wei, Jianxin Wu, Weiyao Lin

In this paper we show that by carefully making good choices for various detailed but important factors in a visual recognition framework using deep learning features, one can achieve a simple, efficient, yet highly accur…

General Classificationimage-classificationImage Classification