paper-with-me

홈 › Papers

Fine-Tuning Whisper for Inclusive Prosodic Stress Analysis

2025-03-03 · Samuel S. Sohn, Sten Knutsen, Karin Stromswold

Prosody plays a crucial role in speech perception, influencing both human understanding and automatic speech recognition (ASR) systems. Despite its importance, prosodic stress remains under-studied due to the challenge of efficiently analyzing it. This study explores fine-tuning OpenAI's Whisper large-v2 ASR model to recognize phrasal, lexical, and contrastive stress in speech. Using a dataset of 66 native English speakers, including male, female, neurotypical, and neurodivergent individuals, we assess the model's ability to generalize stress patterns and classify speakers by neurotype and gender based on brief speech samples. Our results highlight near-human accuracy in ASR performance across all three stress types and near-perfect precision in classifying gender and neurotype. By improving prosody-aware ASR, this work contributes to equitable and robust transcription technologies for diverse populations.

📄 PDF Abstract BibTeX arXiv:2503.02907

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Transformer-based segmentation of prosodic boundaries in Brazilian Portuguese

2026-07-08 · Rodrigo de Freitas Lima, Julio Cesar Galdino, Marcos Vinicius Treviso arxiv

Automatic prosodic segmentation identifies boundaries between speech units from acoustic and linguistic evidence. Although recent deep learning approaches have produced strong results for English, automatic segmentation …

Recovering implicit pitch contours from formants in whispered speech

2023-07-06 · Pablo Pérez Zarazaga, Zofia Malisz

Whispered speech is characterised by a noise-like excitation that results in the lack of fundamental frequency. Considering that prosodic phenomena such as intonation are perceived through f0 variation, the perception of…

Denoising

Multimodal Belief Prediction

2024-06-11 · John Murzaku, Adil Soubki, Owen Rambow

Recognizing a speaker's level of commitment to a belief is a difficult task; humans do not only interpret the meaning of the words in context, but also understand cues from intonation and other aspects of the audio signa…

Prediction

Whispered-to-voiced Alaryngeal Speech Conversion with Generative Adversarial Networks

2018-08-31 · Santiago Pascual, Antonio Bonafonte, Joan Serrà, Jose A. Gonzalez

Most methods of voice restoration for patients suffering from aphonia either produce whispered or monotone speech. Apart from intelligibility, this type of speech lacks expressiveness and naturalness due to the absence o…

Speech EnhancementSpeech Recognition

A Semi-Supervised Framework for Speech Confidence Detection using Whisper

2026-05-12 · Adam Wynn, Jingyun Wang arxiv

Automatic detection of speaker confidence is critical for adaptive computing but remains constrained by limited labelled data and the subjectivity of paralinguistic annotations. This paper proposes a semi-supervised hybr…