paper-with-me

Papers

TitaNet: Neural Model for speaker representation with 1D Depth-wise separable convolutions and global context

2021-10-08 · Nithin Rao Koluguri, Taejin Park, Boris Ginsburg

In this paper, we propose TitaNet, a novel neural network architecture for extracting speaker representations. We employ 1D depth-wise separable convolutions with Squeeze-and-Excitation (SE) layers with global context followed by channel attention based statistics pooling layer to map variable-length utterances to a fixed-length embedding (t-vector). TitaNet is a scalable architecture and achieves state-of-the-art performance on speaker verification task with an equal error rate (EER) of 0.68% on the VoxCeleb1 trial file and also on speaker diarization tasks with diarization error rate (DER) of 1.73% on AMI-MixHeadset, 1.99% on AMI-Lapel and 1.11% on CH109. Furthermore, we investigate various sizes of TitaNet and present a light TitaNet-S model with only 6M parameters that achieve near state-of-the-art results in diarization tasks.

📄 PDF Abstract BibTeX arXiv:2110.04410

Code (2)

NVIDIA/NeMo 공식 구현 pytorch
Wadaboa/titanet pytorch

Tasks

speaker-diarizationSpeaker DiarizationSpeaker Verification

Similar Papers 제목 키워드 기반

A Compact End-to-End Model with Local and Global Context for Spoken Language Identification

2022-10-27 · Fei Jia, Nithin Rao Koluguri, Jagadeesh Balam, Boris Ginsburg

We introduce TitaNet-LID, a compact end-to-end neural network for Spoken Language Identification (LID) that is based on the ContextNet architecture. TitaNet-LID employs 1D depth-wise separable convolutions and Squeeze-an…

Language IdentificationSpoken language identification

Few-Shot Speaker Identification Using Depthwise Separable Convolutional Network with Channel Attention

2022-04-24 · Yanxiong Li, Wucheng Wang, Hao Chen, Wenchang Cao 외

Although few-shot learning has attracted much attention from the fields of image and audio classification, few efforts have been made on few-shot speaker identification. In the task of few-shot learning, overfitting is a…

Audio ClassificationFew-Shot LearningSpeaker Identification

Evaluation of Speech Representations for MOS prediction

2023-06-16 · Frederico S. Oliveira, Edresson Casanova, Arnaldo Cândido Júnior, Lucas R. S. Gris 외

In this paper, we evaluate feature extraction models for predicting speech quality. We also propose a model architecture to compare embeddings of supervised learning and self-supervised learning models with embeddings of…

PredictionSelf-Supervised LearningSpeaker Verification

Depthwise Separable Convolutions for Neural Machine Translation

2017-06-09 · ICLR 2018 1 · Lukasz Kaiser, Aidan N. Gomez, Francois Chollet

Depthwise separable convolutions reduce the number of parameters and computation used in convolutional operations while increasing representational efficiency. They have been shown to be successful in image classificatio…

image-classificationMachine TranslationTranslation

An Integrated Framework for Two-pass Personalized Voice Trigger

2021-06-30 · Dexin Liao, Jing Li, Yiming Zhi, Song Li 외

In this paper, we present the XMUSPEECH system for Task 1 of 2020 Personalized Voice Trigger Challenge (PVTC2020). Task 1 is a joint wake-up word detection with speaker verification on close talking data. The whole syste…

Keyword SpottingMulti-Task LearningSpeaker VerificationVocal Bursts Valence Prediction