paper-with-me

홈 › Papers

Zero-shot Domain-sensitive Speech Recognition with Prompt-conditioning Fine-tuning

2023-07-18 · Feng-Ting Liao, Yung-Chieh Chan, Yi-Chang Chen, Chan-Jan Hsu, Da-Shan Shiu

In this work, we propose a method to create domain-sensitive speech recognition models that utilize textual domain information by conditioning its generation on a given text prompt. This is accomplished by fine-tuning a pre-trained, end-to-end model (Whisper) to learn from demonstrations with prompt examples. We show that this ability can be generalized to different domains and even various prompt contexts, with our model gaining a Word Error Rate (WER) reduction of up to 33% on unseen datasets from various domains, such as medical conversation, air traffic control communication, and financial meetings. Considering the limited availability of audio-transcript pair data, we further extend our method to text-only fine-tuning to achieve domain sensitivity as well as domain adaptation. We demonstrate that our text-only fine-tuned model can also attend to various prompt contexts, with the model reaching the most WER reduction of 29% on the medical conversation dataset.

📄 PDF Abstract BibTeX arXiv:2307.10274

Code (1)

mtkresearch/clairaudience 공식 구현 pytorch

Tasks

Domain Adaptationspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Prompt Amplification and Zero-Shot Late Fusion in Audio-Language Models for Speech Emotion Recognition

2026-03-24 · Saurabh Kataria, Xiao Hu arxiv

Audio-Language Models (ALMs) are making strides in understanding speech and non-speech audio. However, domain-specialist Foundation Models (FMs) remain the best for closed-ended speech processing tasks such as Speech Emo…

Speech Emotion Recognition

Scaling ASR Improves Zero and Few Shot Learning

2021-11-10 · Alex Xiao, Weiyi Zheng, Gil Keren, Duc Le 외

With 4.5 million hours of English speech from 10 different sources across 120 countries and models of up to 10 billion parameters, we explore the frontiers of scale for automatic speech recognition. We propose data selec…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Few-Shot Learningspeech-recognition+1

VoxEmo: Benchmarking Speech Emotion Recognition with Speech LLMs

2026-03-09 · Hezhao Zhang, Huang-Cheng Chou, Shrikanth Narayanan, Thomas Hain arxiv

Speech Large Language Models (LLMs) show great promise for speech emotion recognition (SER) via generative interfaces. However, shifting from closed-set classification to open text generation introduces zero-shot stochas…

Speech Emotion RecognitionText Generation

OpenSR: Open-Modality Speech Recognition via Maintaining Multi-Modality Alignment

2023-06-10 · Xize Cheng, Tao Jin, Linjun Li, Wang Lin 외

Speech Recognition builds a bridge between the multimedia streaming (audio-only, visual-only or audio-visual) and the corresponding text transcription. However, when training the specific model of new domain, it often ge…

Audio-Visual Speech RecognitionLip Readingspeech-recognitionSpeech Recognition+1

Large Language Models Meet Contrastive Learning: Zero-Shot Emotion Recognition Across Languages

2025-03-25 · Heqing Zou, Fengmao Lv, Desheng Zheng, Eng Siong Chng 외

Multilingual speech emotion recognition aims to estimate a speaker's emotional state using a contactless method across different languages. However, variability in voice characteristics and linguistic diversity poses sig…

Contrastive LearningDiversityEmotion RecognitionSpeech Emotion Recognition