paper-with-me

홈 › Papers

VARAN: Variational Inference for Self-Supervised Speech Models Fine-Tuning on Downstream Tasks

2025-08-16 · Daria Diatlova, Nikita Balagansky, Alexander Varlamov, Egor Spirin arxiv

Conventional methods for aggregating layers in fine-tuned self-supervised speech models, such as using the final layer or weighted sum, suffer from information bottlenecks and static feature weighting for all dataset examples. We propose VARAN, a framework that dynamically tailors layer aggregation to individual inputs. By employing layer-specialized probing heads and data-dependent weighting, VARAN adaptively prioritizes layer's features based on input. Evaluations on automatic speech recognition and speech emotion recognition tasks demonstrate VARAN's superior performance, particularly when using the LoRA fine-tuning technique. The framework resolves the trade-off between preserving layer-specific information and enabling flexible feature utilization, advancing efficient adaptation of self-supervised speech representations.

📄 PDF Abstract BibTeX arXiv:2508.12061

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Emotion RecognitionSpeech Recognition

Similar Papers 제목 키워드 기반

A Convolutional Deep Markov Model for Unsupervised Speech Representation Learning

2020-06-03 · Sameer Khurana, Antoine Laurent, Wei-Ning Hsu, Jan Chorowski 외

Probabilistic Latent Variable Models (LVMs) provide an alternative to self-supervised learning approaches for linguistic representation learning from speech. LVMs admit an intuitive probabilistic interpretation where the…

Representation LearningSelf-Supervised LearningSpeech Representation LearningVariational Inference

Posterior sampling algorithms for unsupervised speech enhancement with recurrent variational autoencoder

2023-09-19 · Mostafa Sadeghi, Romain Serizel

In this paper, we address the unsupervised speech enhancement problem based on recurrent variational autoencoder (RVAE). This approach offers promising generalization performance over the supervised counterpart. Neverthe…

Computational EfficiencySpeech EnhancementVariational Inference

Unsupervised TTS Acoustic Modeling for TTS with Conditional Disentangled Sequential VAE

2022-06-06 · Jiachen Lian, Chunlei Zhang, Gopala Krishna Anumanchipalli, Dong Yu

In this paper, we propose a novel unsupervised text-to-speech acoustic model training scheme, named UTTS, which does not require text-audio pairs. UTTS is a multi-speaker speech synthesizer that supports zero-shot voice …

Representation LearningSpeech Representation LearningSpeech Synthesistext-to-speech+2

HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis

2023-11-21 · Sang-Hoon Lee, Ha-Yeong Choi, Seung-bin Kim, Seong-Whan Lee

Large language models (LLM)-based speech synthesis has been widely adopted in zero-shot speech synthesis. However, they require a large-scale data and possess the same limitations as previous autoregressive speech models…

Speech SynthesisSuper-Resolutiontext-to-speechText to Speech+2

Self-Supervised Speech Quality Estimation and Enhancement Using Only Clean Speech

2024-02-26 · Szu-Wei Fu, Kuo-Hsuan Hung, Yu Tsao, Yu-Chiang Frank Wang

Speech quality estimation has recently undergone a paradigm shift from human-hearing expert designs to machine-learning models. However, current models rely mainly on supervised learning, which is time-consuming and expe…

QuantizationSpeech Enhancement