paper-with-me

Papers

A Lightweight Speaker Recognition System Using Timbre Properties

2020-10-12 · Abu Quwsar Ohi, M. F. Mridha, Md. Abdul Hamid, Muhammad Mostafa Monowar, Dongsu Lee, Jinsul Kim

Speaker recognition is an active research area that contains notable usage in biometric security and authentication system. Currently, there exist many well-performing models in the speaker recognition domain. However, most of the advanced models implement deep learning that requires GPU support for real-time speech recognition, and it is not suitable for low-end devices. In this paper, we propose a lightweight text-independent speaker recognition model based on random forest classifier. It also introduces new features that are used for both speaker verification and identification tasks. The proposed model uses human speech based timbral properties as features that are classified using random forest. Timbre refers to the very basic properties of sound that allow listeners to discriminate among them. The prototype uses seven most actively searched timbre properties, boominess, brightness, depth, hardness, roughness, sharpness, and warmth as features of our speaker recognition model. The experiment is carried out on speaker verification and speaker identification tasks and shows the achievements and drawbacks of the proposed model. In the speaker identification phase, it achieves a maximum accuracy of 78%. On the contrary, in the speaker verification phase, the model maintains an accuracy of 80% having an equal error rate (ERR) of 0.24.

📄 PDF Abstract BibTeX arXiv:2010.05502

Code (0)

등록된 구현이 없습니다.

Tasks

GPUSpeaker IdentificationSpeaker RecognitionSpeaker Verificationspeech-recognitionSpeech RecognitionText-Independent Speaker Recognition

Similar Papers 제목 키워드 기반

Streaming Voice Conversion Via Intermediate Bottleneck Features And Non-streaming Teacher Guidance

2022-10-27 · Yuanzhe Chen, Ming Tu, Tang Li, Xin Li 외

Streaming voice conversion (VC) is the task of converting the voice of one person to another in real-time. Previous streaming VC methods use phonetic posteriorgrams (PPGs) extracted from automatic speech recognition (ASR…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

CTEFM-VC: Zero-Shot Voice Conversion Based on Content-Aware Timbre Ensemble Modeling and Flow Matching

2024-11-04 · Yu Pan, Yuguang Yang, Jixun Yao, Jianhao Ye 외

Zero-shot voice conversion (VC) aims to transform the timbre of a source speaker into any previously unseen target speaker, while preserving the original linguistic content. Despite notable progress, attaining a degree o…

Speaker VerificationVoice Conversion

JukeBox: A Multilingual Singer Recognition Dataset

2020-08-08 · Anurag Chowdhury, Austin Cozzo, Arun Ross

A text-independent speaker recognition system relies on successfully encoding speech factors such as vocal pitch, intensity, and timbre to achieve good performance. A majority of such systems are trained and evaluated us…

Speaker RecognitionText-Independent Speaker Recognition

StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching

2024-12-06 · Jixun Yao, Yuguang Yang, Yu Pan, Ziqian Ning 외

Zero-shot voice conversion (VC) aims to transfer the timbre from the source speaker to an arbitrary unseen speaker while preserving the original linguistic content. Despite recent advancements in zero-shot VC using langu…

Voice Conversion

VITS-Based Singing Voice Conversion Leveraging Whisper and multi-scale F0 Modeling

2023-10-04 · Ziqian Ning, Yuepeng Jiang, Zhichao Wang, Bin Zhang 외

This paper introduces the T23 team's system submitted to the Singing Voice Conversion Challenge 2023. Following the recognition-synthesis framework, our singing conversion model is based on VITS, incorporating four key m…

DecoderVoice Conversion