VARAN: Variational Inference for Self-Supervised Speech Models Fine-Tuning on Downstream Tasks
Conventional methods for aggregating layers in fine-tuned self-supervised speech models, such as using the final layer or weighted sum, suffer from information bottlenecks and static feature weighting for all dataset examples. We propose VARAN, a framework that dynamically tailors layer aggregation to individual inputs. By employing layer-specialized probing heads and data-dependent weighting, VARAN adaptively prioritizes layer's features based on input. Evaluations on automatic speech recognition and speech emotion recognition tasks demonstrate VARAN's superior performance, particularly when using the LoRA fine-tuning technique. The framework resolves the trade-off between preserving layer-specific information and enabling flexible feature utilization, advancing efficient adaptation of self-supervised speech representations.
Code (0)
등록된 구현이 없습니다.
Tasks
Speech Emotion RecognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
A Convolutional Deep Markov Model for Unsupervised Speech Representation Learning
Probabilistic Latent Variable Models (LVMs) provide an alternative to self-supervised learning approaches for linguistic representation learning from speech. LVMs admit an intuitive probabilistic interpretation where the…
Representation LearningSelf-Supervised LearningSpeech Representation LearningVariational InferencePosterior sampling algorithms for unsupervised speech enhancement with recurrent variational autoencoder
In this paper, we address the unsupervised speech enhancement problem based on recurrent variational autoencoder (RVAE). This approach offers promising generalization performance over the supervised counterpart. Neverthe…
Computational EfficiencySpeech EnhancementVariational InferenceUnsupervised TTS Acoustic Modeling for TTS with Conditional Disentangled Sequential VAE
In this paper, we propose a novel unsupervised text-to-speech acoustic model training scheme, named UTTS, which does not require text-audio pairs. UTTS is a multi-speaker speech synthesizer that supports zero-shot voice …
Representation LearningSpeech Representation LearningSpeech Synthesistext-to-speech+2HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis
Large language models (LLM)-based speech synthesis has been widely adopted in zero-shot speech synthesis. However, they require a large-scale data and possess the same limitations as previous autoregressive speech models…
Speech SynthesisSuper-Resolutiontext-to-speechText to Speech+2Self-Supervised Speech Quality Estimation and Enhancement Using Only Clean Speech
Speech quality estimation has recently undergone a paradigm shift from human-hearing expert designs to machine-learning models. However, current models rely mainly on supervised learning, which is time-consuming and expe…
QuantizationSpeech Enhancement