paper-with-me

Papers

Voice Signal Processing for Machine Learning. The Case of Speaker Isolation

2024-03-29 · Radan Ganchev

The widespread use of automated voice assistants along with other recent technological developments have increased the demand for applications that process audio signals and human voice in particular. Voice recognition tasks are typically performed using artificial intelligence and machine learning models. Even though end-to-end models exist, properly pre-processing the signal can greatly reduce the complexity of the task and allow it to be solved with a simpler ML model and fewer computational resources. However, ML engineers who work on such tasks might not have a background in signal processing which is an entirely different area of expertise. The objective of this work is to provide a concise comparative analysis of Fourier and Wavelet transforms that are most commonly used as signal decomposition methods for audio processing tasks. Metrics for evaluating speech intelligibility are also discussed, namely Scale-Invariant Signal-to-Distortion Ratio (SI-SDR), Perceptual Evaluation of Speech Quality (PESQ), and Short-Time Objective Intelligibility (STOI). The level of detail in the exposition is meant to be sufficient for an ML engineer to make informed decisions when choosing, fine-tuning, and evaluating a decomposition method for a specific ML model. The exposition contains mathematical definitions of the relevant concepts accompanied with intuitive non-mathematical explanations in order to make the text more accessible to engineers without deep expertise in signal processing. Formal mathematical definitions and proofs of theorems are intentionally omitted in order to keep the text concise.

📄 PDF Abstract BibTeX arXiv:2403.20202

Code (1)

rganchev/speech-signal-processing-for-ml 공식 구현

Similar Papers 제목 키워드 기반

Asynchronous Voice Anonymization Using Adversarial Perturbation On Speaker Embedding

2024-06-12 · Rui Wang, Liping Chen, Kong Aik Lee, Zhen-Hua Ling

Voice anonymization has been developed as a technique for preserving privacy by replacing the speaker's voice in a speech signal with that of a pseudo-speaker, thereby obscuring the original voice attributes from machine…

Disentanglement

Practical Hidden Voice Attacks against Speech and Speaker Recognition Systems

2019-03-18 · Hadi Abdullah, Washington Garcia, Christian Peeters, Patrick Traynor 외

Voice Processing Systems (VPSes), now widely deployed, have been made significantly more accurate through the application of recent advances in machine learning. However, adversarial machine learning has similarly advanc…

Audio Signal ProcessingBIG-bench Machine LearningSpeaker Recognition

ALO-VC: Any-to-any Low-latency One-shot Voice Conversion

2023-06-01 · Bohan Wang, Damien Ronssin, Milos Cernak

This paper presents ALO-VC, a non-parallel low-latency one-shot phonetic posteriorgrams (PPGs) based voice conversion method. ALO-VC enables any-to-any voice conversion using only one utterance from the target speaker, w…

CPUVoice Conversion

VoiceFilter: Targeted Voice Separation by Speaker-Conditioned Spectrogram Masking

2018-10-11 · Quan Wang, Hannah Muckenhirn, Kevin Wilson, Prashant Sridhar 외

In this paper, we present a novel system that separates the voice of a target speaker from multi-speaker signals, by making use of a reference signal from the target speaker. We achieve this by training two separate neur…

Speaker RecognitionSpeaker SeparationSpeech Enhancementspeech-recognition+1

Attacking Voice Anonymization Systems with Augmented Feature and Speaker Identity Difference

2024-12-26 · Yanzhe Zhang, Zhonghao Bi, Feiyang Xiao, Xuefeng Yang 외

This study focuses on the First VoicePrivacy Attacker Challenge within the ICASSP 2025 Signal Processing Grand Challenge, which aims to develop speaker verification systems capable of determining whether two anonymized s…

Data AugmentationSpeaker Verification