paper-with-me

홈 › Papers

Evaluating the Effectiveness of Transformer Layers in Wav2Vec 2.0, XLS-R, and Whisper for Speaker Identification Tasks

2025-08-29 · Linus Stuhlmann, Michael Alexander Saxer arxiv

This study evaluates the performance of three advanced speech encoder models, Wav2Vec 2.0, XLS-R, and Whisper, in speaker identification tasks. By fine-tuning these models and analyzing their layer-wise representations using SVCCA, k-means clustering, and t-SNE visualizations, we found that Wav2Vec 2.0 and XLS-R capture speaker-specific features effectively in their early layers, with fine-tuning improving stability and performance. Whisper showed better performance in deeper layers. Additionally, we determined the optimal number of transformer layers for each model when fine-tuned for speaker identification tasks.

📄 PDF Abstract BibTeX arXiv:2509.00230

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker Identification

Similar Papers 제목 키워드 기반

BaldWhisper: Faster Whisper with Head Shearing and Layer Merging

2025-10-06 · Yaya Sy, Christophe Cerisara, Irina Illina arxiv

Pruning large pre-trained transformers in a data-scarce scenario is challenging, as it often requires massive retraining data to recover performance. For instance, Distill-Whisper prunes Whisper by 40 and retrains on 21,…

Target Speaker ASR with Whisper

2024-09-14 · Alexander Polok, Dominik Klement, Matthew Wiesner, Sanjeev Khudanpur 외

We propose a novel approach to enable the use of large, single-speaker ASR models, such as Whisper, for target speaker ASR. The key claim of this method is that it is much easier to model relative differences among speak…

Speech Separation

The Voice of Equity: A Systematic Evaluation of Bias Mitigation Techniques for Speech-Based Cognitive Impairment Detection Across Architectures and Demographics

2026-01-07 · Yasaman Haghbin, Sina Rashidi, Ali Zolnour, Maryam Zolnoori arxiv

Speech-based detection of cognitive impairment offers a scalable, non-invasive screening, yet algorithmic bias across demographic and linguistic subgroups remains critically underexplored. We present the first comprehens…

Whisper Speaker Identification: Leveraging Pre-Trained Multilingual Transformers for Robust Speaker Embeddings

2025-03-13 · Jakaria Islam Emon, Md Abu Salek, Kazi Tamanna Alam

Speaker identification in multilingual settings presents unique challenges, particularly when conventional models are predominantly trained on English data. In this paper, we propose WSI (Whisper Speaker Identification),…

Speaker Identificationspeech-recognitionSpeech Recognition

AdaPTwin: Low-Cost Adaptive Compression of Product Twins in Transformers

2024-06-13 · Emil Biju, Anirudh Sriram, Mert Pilanci

While large transformer-based models have exhibited remarkable performance in speaker-independent speech recognition, their large size and computational requirements make them expensive or impractical to use in resource-…

speech-recognitionSpeech Recognition