paper-with-me

홈 › Papers

Layer-wise Probing of wav2vec 2.0 and Whisper for Consonant Cluster Reduction in African American English

2026-06-22 · Hamid Mojarad, Kevin Tang arxiv

Self-supervised and supervised speech models are increasingly used to investigate which linguistic information their internal representations encode, and at what level of abstraction they encode it. One underexplored phenomenon is consonant cluster reduction (CCR) in African American English (AAE), a widespread phonological process and a source of automatic speech recognition (ASR) disparity. To examine how CCR is represented, we conduct speaker-independent layer-wise probing of wav2vec2-base and Whisper-small using two tasks: segmental reduction detection and segmental restoration of underlying cluster identity. Both models distinguish reduced and canonical forms with high accuracy. Crucially, reduced segments retain cues to their underlying stops, indicating that CCR is encoded as structured gradient phonological variation rather than simple segmental deletion. These results demonstrate structured phonological encoding of AAE CCR patterns in modern speech models.

📄 PDF Abstract BibTeX arXiv:2606.23948

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Recognition

Similar Papers 제목 키워드 기반

Fine-tuning Whisper for Pashto ASR: strategies and scale

2026-04-07 · Hanif Rahman arxiv

Pashto is absent from Whisper's pre-training corpus despite being one of CommonVoice's largest language collections, leaving off-the-shelf models unusable: all Whisper sizes output Arabic, Dari, or Urdu script on Pashto …

Evaluating the Effectiveness of Transformer Layers in Wav2Vec 2.0, XLS-R, and Whisper for Speaker Identification Tasks

2025-08-29 · Linus Stuhlmann, Michael Alexander Saxer arxiv

This study evaluates the performance of three advanced speech encoder models, Wav2Vec 2.0, XLS-R, and Whisper, in speaker identification tasks. By fine-tuning these models and analyzing their layer-wise representations u…

Speaker Identification

From Text Metrics to Model Internals: A Study of Whisper ASR Hallucination Detection

2026-06-22 · Jan Jasiński, Mateusz Barański, Julitta Bartolewska, Marcin Witkowski 외 arxiv

Hallucinations of ASR models - fluent transcriptions with no basis in audio - degrade system performance and pose risks in downstream applications. Robust detection of such errors remains a challenge. This paper studies …

Understanding Multilingual Medical ASR Adaptation Through Layer-Wise Analysis

2026-08-19 · Souranil Kahali, Rituparna Bose, Abner Hernandez, Tomas Arias-Vergara 외 arxiv

Medical automatic speech recognition (MedASR) requires adaptation to specialised terminology, limited annotated clinical data, and multilingual use cases. Although large-scale pretrained ASR models such as Whisper achiev…

Speech Recognition

Beyond WER: Probing Whisper's Sub-token Decoder Across Diverse Language Resource Levels

2025-09-29 · Siyu Liang, Nicolas Ballier, Gina-Anne Levow, Richard Wright arxiv

While large multilingual automatic speech recognition (ASR) models achieve remarkable performance, the internal mechanisms of the end-to-end pipeline, particularly concerning fairness and efficacy across languages, remai…

Speech Recognition