paper-with-me

Papers

Investigating self-supervised front ends for speech spoofing countermeasures

2021-11-15 · Xin Wang, Junichi Yamagishi

Self-supervised speech model is a rapid progressing research topic, and many pre-trained models have been released and used in various down stream tasks. For speech anti-spoofing, most countermeasures (CMs) use signal processing algorithms to extract acoustic features for classification. In this study, we use pre-trained self-supervised speech models as the front end of spoofing CMs. We investigated different back end architectures to be combined with the self-supervised front end, the effectiveness of fine-tuning the front end, and the performance of using different pre-trained self-supervised models. Our findings showed that, when a good pre-trained front end was fine-tuned with either a shallow or a deep neural network-based back end on the ASVspoof 2019 logical access (LA) training set, the resulting CM not only achieved a low EER score on the 2019 LA test set but also significantly outperformed the baseline on the ASVspoof 2015, 2021 LA, and 2021 deepfake test sets. A sub-band analysis further demonstrated that the CM mainly used the information in a specific frequency band to discriminate the bona fide and spoofed trials across the test sets.

📄 PDF Abstract BibTeX arXiv:2111.07725

Code (1)

asvspoof-challenge/2021 공식 구현 pytorch

Tasks

Face Swapping

Similar Papers 제목 키워드 기반

Front-End Adapter: Adapting Front-End Input of Speech based Self-Supervised Learning for Speech Recognition

2023-02-18 · Xie Chen, Ziyang Ma, Changli Tang, Yujin Wang 외

Recent years have witnessed a boom in self-supervised learning (SSL) in various areas including speech processing. Speech based SSL models present promising performance in a range of speech related tasks. However, the tr…

Self-Supervised Learningspeech-recognitionSpeech Recognition

Frontend Token Enhancement for Token-Based Speech Recognition

2026-02-04 · Takanori Ashihara, Shota Horiguchi, Kohei Matsuura, Tsubasa Ochiai 외 arxiv

Discretized representations of speech signals are efficient alternatives to continuous features for various speech applications, including automatic speech recognition (ASR) and speech language models. However, these rep…

Self-Supervised LearningSpeech Recognition

EURO: ESPnet Unsupervised ASR Open-source Toolkit

2022-11-30 · Dongji Gao, Jiatong Shi, Shun-Po Chuang, Leibny Paola Garcia 외

This paper describes the ESPnet Unsupervised ASR Open-source Toolkit (EURO), an end-to-end open-source toolkit for unsupervised automatic speech recognition (UASR). EURO adopts the state-of-the-art UASR learning method i…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Leveraging Unimodal Self-Supervised Learning for Multimodal Audio-Visual Speech Recognition

2022-02-24 · ACL 2022 5 · Xichen Pan, Peiyu Chen, Yichen Gong, Helong Zhou 외

Training Transformer-based models demands a large amount of data, while obtaining aligned and labelled data in multimodality is rather cost-demanding, especially for audio-visual speech recognition (AVSR). Thus it makes …

Audio-Visual Speech RecognitionAutomatic Speech Recognition (ASR)Language ModellingLipreading+6

Two Front-Ends, One Model : Fusing Heterogeneous Speech Features for Low Resource ASR with Multilingual Pre-Training

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Transfer learning is widely applied in various deep learning-based speech tasks, especially for tasks with a limited amount of data. Recent studies in transfer learning mainly focused on either supervised or self-supervi…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1