paper-with-me

홈 › Papers

A Pre-training Framework that Encodes Noise Information for Speech Quality Assessment

2024-11-07 · Subrina Sultana, Donald S. Williamson

Self-supervised learning (SSL) has grown in interest within the speech processing community, since it produces representations that are useful for many downstream tasks. SSL uses global and contextual methods to produce robust representations, where SSL even outperforms supervised models. Most self-supervised approaches, however, are limited to embedding information about, i.e., the phonemes, speaker identity, and emotion, into the extracted representations, where they become invariant to background sounds due to contrastive and auto-regressive learning. This is limiting because many downstream tasks leverage noise information to function accurately. Therefore, we propose a pre-training framework that learns information pertaining to background noise in a supervised manner, while jointly embedding speech information using a self-supervised strategy. We experiment with multiple encoders and show that our framework is useful for perceptual speech quality estimation, which relies on background cues. Our results show that the proposed approach improves performance with fewer parameters, in comparison to multiple baselines.

📄 PDF Abstract BibTeX arXiv:2411.04379

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised Learning

Similar Papers 제목 키워드 기반

Mutual Information Maximization for Simple and Accurate Part-Of-Speech Induction

2018-04-20 · NAACL 2019 6 · Karl Stratos

We address part-of-speech (POS) induction by maximizing the mutual information between the induced label and its context. We focus on two training objectives that are amenable to stochastic gradient descent (SGD): a nove…

ClusteringPOS

Disentangling speech from surroundings with neural embeddings

2022-03-29 · Ahmed Omran, Neil Zeghidour, Zalán Borsos, Félix de Chaumont Quitry 외

We present a method to separate speech signals from noisy environments in the embedding space of a neural audio codec. We introduce a new training procedure that allows our model to produce structured encodings of audio …

Attribute

Lombard Effect for Bilingual Speakers in Cantonese and English: importance of spectro-temporal features

2022-04-14 · Maximilian Karl Scharf, Sabine Hochmuth, Lena L. N. Wong, Birger Kollmeier 외

For a better understanding of the mechanisms underlying speech perception and the contribution of different signal features, computational models of speech recognition have a long tradition in hearing research. Due to th…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Unsupervised Latent Behavior Manifold Learning from Acoustic Features: audio2behavior

2017-01-12 · Haoqi Li, Brian Baucom, Panayiotis Georgiou

Behavioral annotation using signal processing and machine learning is highly dependent on training data and manual annotations of behavioral labels. Previous studies have shown that speech information encodes significant…

Interpreting Speaker Characteristics in the Dimensions of Self-Supervised Speech Features

2026-03-03 · Kyle Janse van Rensburg, Benjamin van Niekerk, Herman Kamper arxiv

How do speech models trained through self-supervised learning structure their representations? Previous studies have looked at how information is encoded in feature vectors across different layers. But few studies have c…

Self-Supervised Learning