paper-with-me

홈 › Papers

A Unified Deep Neural Network for Speaker and Language Recognition

2015-04-03 · Fred Richardson, Douglas Reynolds, Najim Dehak

Learned feature representations and sub-phoneme posteriors from Deep Neural Networks (DNNs) have been used separately to produce significant performance gains for speaker and language recognition tasks. In this work we show how these gains are possible using a single DNN for both speaker and language recognition. The unified DNN approach is shown to yield substantial performance improvements on the the 2013 Domain Adaptation Challenge speaker recognition task (55% reduction in EER for the out-of-domain condition) and on the NIST 2011 Language Recognition Evaluation (48% reduction in EER for the 30s test condition).

📄 PDF Abstract BibTeX arXiv:1504.00923

Code (0)

등록된 구현이 없습니다.

Tasks

Domain AdaptationSpeaker Recognition

Similar Papers 제목 키워드 기반

Collaborative Learning for Language and Speaker Recognition

2016-09-27 · Lantian Li, Zhiyuan Tang, Dong Wang, Andrew Abel 외

This paper presents a unified model to perform language and speaker recognition simultaneously and altogether. The model is based on a multi-task recurrent neural network where the output of one task is fed as the input …

Speaker Recognition

Multi-task Recurrent Model for Speech and Speaker Recognition

2016-03-31 · Zhiyuan Tang, Lantian Li, Dong Wang

Although highly correlated, speech and speaker recognition have been regarded as two independent tasks and studied by two communities. This is certainly not the way that people behave: we decipher both speech content and…

Speaker Recognition

Exploring the Encoding Layer and Loss Function in End-to-End Speaker and Language Recognition System

2018-04-14 · Weicheng Cai, Jinkun Chen, Ming Li

In this paper, we explore the encoding/pooling layer and loss function in the end-to-end speaker and language recognition system. First, a unified and interpretable end-to-end system for both speaker and language recogni…

Speaker Verification

SpeakerLM: End-to-End Versatile Speaker Diarization and Recognition with Multimodal Large Language Models

2025-08-08 · Han Yin, Yafeng Chen, Chong Deng, Luyao Cheng 외 arxiv

The Speaker Diarization and Recognition (SDR) task aims to predict "who spoke when and what" within an audio clip, which is a crucial task in various real-world multi-speaker scenarios such as meeting transcription and d…

Speaker DiarizationSpeech Recognition

Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models

2025-06-06 · Yuke Lin, Ming Cheng, Ze Li, Beilong Tang 외

Multi-speaker automatic speech recognition (MS-ASR) faces significant challenges in transcribing overlapped speech, a task critical for applications like meeting transcription and conversational analysis. While serialize…

Automatic Speech Recognitionspeaker-diarizationSpeaker Diarizationspeech-recognition+1