paper-with-me

홈 › Papers

Sounding Like a Winner? Prosodic Differences in Post-Match Interviews

2025-06-02 · Sofoklis Kakouros, Haoyu Chen

This study examines the prosodic characteristics associated with winning and losing in post-match tennis interviews. Additionally, this research explores the potential to classify match outcomes solely based on post-match interview recordings using prosodic features and self-supervised learning (SSL) representations. By analyzing prosodic elements such as pitch and intensity, alongside SSL models like Wav2Vec 2.0 and HuBERT, the aim is to determine whether an athlete has won or lost their match. Traditional acoustic features and deep speech representations are extracted from the data, and machine learning classifiers are employed to distinguish between winning and losing players. Results indicate that SSL representations effectively differentiate between winning and losing outcomes, capturing subtle speech patterns linked to emotional states. At the same time, prosodic cues -- such as pitch variability -- remain strong indicators of victory.

📄 PDF Abstract BibTeX arXiv:2506.02283

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised Learning

Similar Papers 제목 키워드 기반

Prosody of speech production in latent post-stroke aphasia

2024-08-21 · Cong Zhang, Tong Li, Gayle DeDe, Christos Salis

This study explores prosodic production in latent aphasia, a mild form of aphasia associated with left-hemisphere brain damage (e.g. stroke). Unlike prior research on moderate to severe aphasia, we investigated latent ap…

Diagnostic

The Prosody of Emojis

2025-08-01 · Giulio Zhou, Tsz Kin Lam, Alexandra Birch, Barry Haddow arxiv

Prosodic features such as pitch, timing, and intonation are central to spoken communication, conveying emotion, intent, and discourse structure. In text-based settings, where these cues are absent, emojis act as visual s…

Improving TTS for Shanghainese: Addressing Tone Sandhi via Word Segmentation

2023-07-30 · Yuanhao Chen

Tone is a crucial component of the prosody of Shanghainese, a Wu Chinese variety spoken primarily in urban Shanghai. Tone sandhi, which applies to all multi-syllabic words in Shanghainese, then, is key to natural-soundin…

text-to-speechText to Speech

We Need Variations in Speech Generation: Sub-center Modelling for Speaker Embeddings

2024-07-05 · Ismail Rasim Ulgen, Carlos Busso, John H. L. Hansen, Berrak Sisman

Modeling the rich prosodic variations inherent in human speech is essential for generating natural-sounding speech. While speaker embeddings are commonly used as conditioning inputs in personalized speech generation, the…

Speaker RecognitionSpeech SynthesisVoice Conversion

Towards cross-language prosody transfer for dialog

2023-07-09 · Jonathan E. Avila, Nigel G. Ward

Speech-to-speech translation systems today do not adequately support use for dialog purposes. In particular, nuances of speaker intent and stance can be lost due to improper prosody transfer. We present an exploration of…

Speech-to-Speech TranslationTranslation