paper-with-me

홈 › Papers

Linear Script Representations in Speech Foundation Models Enable Zero-Shot Transliteration

2026-01-06 · Ryan Soh-Eun Shim, Kwanghee Choi, Kalvin Chang, Ming-Hao Hsu, Florian Eichin, Zhizheng Wu, Alane Suhr, Michael A. Hedderich, David Harwath, David R. Mortensen, Barbara Plank arxiv

Multilingual speech foundation models such as Whisper are trained on web-scale data, where data for each language consists of a myriad of regional varieties. However, different regional varieties often employ different scripts to write the same language, rendering speech recognition output also subject to non-determinism in the output script. To mitigate this problem, we show that script is linearly encoded in the activation space of multilingual speech models, and that modifying activations at inference time enables direct control over output script. We find the addition of such script vectors to activations at test time can induce script change even in unconventional language-script pairings (e.g. Italian in Cyrillic and Japanese in Latin script). We apply this approach to inducing post-hoc control over the script of speech recognition output, where we observe competitive performance across all model sizes of Whisper.

📄 PDF Abstract BibTeX arXiv:2601.02906

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Recognition

Similar Papers 제목 키워드 기반

Self-supervised representations in speech-based depression detection

2023-05-20 · Wen Wu, Chao Zhang, Philip C. Woodland

This paper proposes handling training data sparsity in speech-based automatic depression detection (SDD) using foundation models pre-trained with self-supervised learning (SSL). An analysis of SSL representations derived…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Depression DetectionEmotion Recognition+4

Controlling Whisper: Universal Acoustic Adversarial Attacks to Control Speech Foundation Models

2024-07-05 · Vyas Raina, Mark Gales

Speech enabled foundation models, either in the form of flexible speech recognition based systems or audio-prompted large language models (LLMs), are becoming increasingly popular. One of the interesting aspects of these…

Adversarial AttackAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Form+3

Learning Multiple Utterance-Level Attribute Representations with a Unified Speech Encoder

2026-03-09 · Maryem Bouziane, Salima Mdhaffar, Yannick Estève arxiv

Speech foundation models trained with self-supervised learning produce generic speech representations that support a wide range of speech processing tasks. When further adapted with supervised learning, these models can …

Self-Supervised LearningSpeaker Recognition

Impact of automatic speech recognition quality on Alzheimer's disease detection from spontaneous speech: a reproducible benchmark study with lexical modeling and statistical validation

2026-03-18 · Himadri S Samanta arxiv

Early detection of Alzheimer's disease from spontaneous speech has emerged as a promising non-invasive screening approach. However, the influence of automatic speech recognition (ASR) quality on downstream clinical langu…

Alzheimer's Disease DetectionSpeech Recognition

Unsupervised low-rank representations for speech emotion recognition

2021-04-14 · Georgios Paraskevopoulos, Efthymios Tzinis, Nikolaos Ellinas, Theodoros Giannakopoulos 외

We examine the use of linear and non-linear dimensionality reduction algorithms for extracting low-rank feature representations for speech emotion recognition. Two feature sets are used, one based on low-level descriptor…

Dimensionality ReductionEmotion RecognitionGeneral ClassificationSpeech Emotion Recognition