paper-with-me

Papers

WhisTLE: Deeply Supervised, Text-Only Domain Adaptation for Pretrained Speech Recognition Transformers

2025-09-12 · Akshat Pandey, Karun Kumar, Raphael Tang arxiv

Pretrained automatic speech recognition (ASR) models such as Whisper perform well but still need domain adaptation to handle unseen parlance. In many real-world settings, collecting speech data is impractical, necessitating text-only adaptation. We propose WhisTLE, a deeply supervised, text-only adaptation method for pretrained encoder-decoder ASR models. WhisTLE trains a variational autoencoder (VAE) to model encoder outputs from text and fine-tunes the decoder using the learned text-to-latent encoder, optionally combined with text-to-speech (TTS) adaptation. At inference, the original encoder is restored, incurring no extra runtime cost. Across four datasets and four ASR models, WhisTLE with TTS reduces word error rate (WER) by a relative 49.0% and outperforms all non-WhisTLE baselines in 100 of 112 scenarios. We also find that WhisTLE additively complements any combination of other domain adaptation approaches; we thus recommend the inclusion of WhisTLE during standard processes for adapting encoder-decoder ASR models.

📄 PDF Abstract BibTeX arXiv:2509.10452

Code (0)

등록된 구현이 없습니다.

Tasks

Speech RecognitionDomain Adaptation

Similar Papers 제목 키워드 기반

DNA: Deeply-supervised Nonlinear Aggregation for Salient Object Detection

2019-03-28 · Yun Liu, Ming-Ming Cheng, Xin-Yu Zhang, Guang-Yu Nie 외

Recent progress on salient object detection mainly aims at exploiting how to effectively integrate multi-scale convolutional features in convolutional neural networks (CNNs). Many popular methods impose deep supervision …

object-detectionObject DetectionRGB Salient Object DetectionSaliency Prediction+1

Dolph2Vec: Self-Supervised Representations of Dolphin Vocalizations

2026-06-10 · Chiara Semenzin, Faadil Mustun, Roberto Dessi, Pierre Orhan 외 arxiv

Self-supervised learning (SSL) has opened new opportunities in bioacoustics by enabling scalable modeling of animal vocalizations without the need for expensive manual annotation. However, current SSL models in this doma…

Self-Supervised Learning

MECPformer: Multi-estimations Complementary Patch with CNN-Transformers for Weakly Supervised Semantic Segmentation

2023-03-19 · Chunmeng Liu, Guangyao Li, Yao Shen, Ruiqi Wang

The initial seed based on the convolutional neural network (CNN) for weakly supervised semantic segmentation always highlights the most discriminative regions but fails to identify the global target information. Methods …

Semantic SegmentationWeakly supervised Semantic SegmentationWeakly-Supervised Semantic Segmentation

Deeply-Supervised Recurrent Convolutional Neural Network for Saliency Detection

2016-08-18 · Youbao Tang, Xiangqian Wu, Wei Bu

This paper proposes a novel saliency detection method by developing a deeply-supervised recurrent convolutional neural network (DSRCNN), which performs a full image-to-image saliency prediction. For saliency detection, t…

Saliency DetectionSaliency Prediction

Learning Comment Controversy Prediction in Web Discussions Using Incidentally Supervised Multi-Task CNNs

2018-10-01 · WS 2018 10 · Nils Rethmeier, Marc H{\"u}bner, Leonhard Hennig

Comments on web news contain controversies that manifest as inter-group agreement-conflicts. Tracking such \textit{rapidly evolving controversy} could ease conflict resolution or journalist-user interaction. However, thi…

Binary ClassificationLanguage ModelingLanguage Modelling