paper-with-me

홈 › Papers

Cross-domain Voice Activity Detection with Self-Supervised Representations

2022-09-22 · Sina Alisamir, Fabien Ringeval, Francois Portet

Voice Activity Detection (VAD) aims at detecting speech segments on an audio signal, which is a necessary first step for many today's speech based applications. Current state-of-the-art methods focus on training a neural network exploiting features directly contained in the acoustics, such as Mel Filter Banks (MFBs). Such methods therefore require an extra normalisation step to adapt to a new domain where the acoustics is impacted, which can be simply due to a change of speaker, microphone, or environment. In addition, this normalisation step is usually a rather rudimentary method that has certain limitations, such as being highly susceptible to the amount of data available for the new domain. Here, we exploited the crowd-sourced Common Voice (CV) corpus to show that representations based on Self-Supervised Learning (SSL) can adapt well to different domains, because they are computed with contextualised representations of speech across multiple domains. SSL representations also achieve better results than systems based on hand-crafted representations (MFBs), and off-the-shelf VADs, with significant improvement in cross-domain settings.

📄 PDF Abstract BibTeX arXiv:2209.11061

Code (0)

등록된 구현이 없습니다.

Tasks

Action DetectionActivity DetectionSelf-Supervised Learning

Similar Papers 제목 키워드 기반

Multitask Detection of Speaker Changes, Overlapping Speech and Voice Activity Using wav2vec 2.0

2022-10-26 · Marie Kunešová, Zbyněk Zajíc

Self-supervised learning approaches have lately achieved great success on a broad spectrum of machine learning problems. In the field of speech processing, one of the most successful recent self-supervised models is wav2…

Action DetectionActivity DetectionChange DetectionSelf-Supervised Learning

Self-Adaptive Soft Voice Activity Detection using Deep Neural Networks for Robust Speaker Verification

2019-09-26 · Youngmoon Jung, Yeunju Choi, Hoirin Kim

Voice activity detection (VAD), which classifies frames as speech or non-speech, is an important module in many speech applications including speaker verification. In this paper, we propose a novel method, called self-ad…

Action DetectionActivity DetectionDomain AdaptationSpeaker Verification+1

An End-to-End Architecture for Keyword Spotting and Voice Activity Detection

2016-11-28 · Chris Lengerich, Awni Hannun

We propose a single neural network architecture for two tasks: on-line keyword spotting and voice activity detection. We develop novel inference algorithms for an end-to-end Recurrent Neural Network trained with the Conn…

Action DetectionActivity DetectionGeneral ClassificationKeyword Spotting

Self-supervised Pretraining for Robust Personalized Voice Activity Detection in Adverse Conditions

2023-12-27 · Holger Severin Bovbjerg, Jesper Jensen, Jan Østergaard, Zheng-Hua Tan

In this paper, we propose the use of self-supervised pretraining on a large unlabelled data set to improve the performance of a personalized voice activity detection (VAD) model in adverse conditions. We pretrain a long …

Action DetectionActivity DetectionDenoising

Cross-Channel Attention-Based Target Speaker Voice Activity Detection: Experimental Results for M2MeT Challenge

2022-02-06 · Weiqing Wang, Xiaoyi Qin, Ming Li

In this paper, we present the speaker diarization system for the Multi-channel Multi-party Meeting Transcription Challenge (M2MeT) from team DKU_DukeECE. As the highly overlapped speech exists in the dataset, we employ a…

Action DetectionActivity Detectionspeaker-diarizationSpeaker Diarization