paper-with-me

Papers

Voice Quality Dimensions as Interpretable Primitives for Speaking Style for Atypical Speech and Affect

2025-05-27 · Jaya Narain, Vasudha Kowtha, Colin Lea, Lauren Tooley, Dianna Yee, Vikramjit Mitra, Zifang Huang, Miquel Espi Marques, Jon Huang, Carlos Avendano, Shirley Ren

Perceptual voice quality dimensions describe key characteristics of atypical speech and other speech modulations. Here we develop and evaluate voice quality models for seven voice and speech dimensions (intelligibility, imprecise consonants, harsh voice, naturalness, monoloudness, monopitch, and breathiness). Probes were trained on the public Speech Accessibility (SAP) project dataset with 11,184 samples from 434 speakers, using embeddings from frozen pre-trained models as features. We found that our probes had both strong performance and strong generalization across speech elicitation categories in the SAP dataset. We further validated zero-shot performance on additional datasets, encompassing unseen languages and tasks: Italian atypical speech, English atypical speech, and affective speech. The strong zero-shot performance and the interpretability of results across an array of evaluations suggests the utility of using voice quality dimensions in speaking style-related tasks.

📄 PDF Abstract BibTeX arXiv:2505.21809

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

StepAudio 3 Realtime Technical Report

2026-09-12 · Bin Lin, Bo Zhao, Boyang Zhang, Boyong Wu 외 hf

Realtime spoken interaction demands deep reasoning, prompt responses, and fluid turn-taking. We present StepAudio 3 Realtime, an audio-language foundation model organized around a continuous listen-converse-think-act loo…

VoiceCoach: Interactive Evidence-based Training for Voice Modulation Skills in Public Speaking

2020-01-22 · Xingbo Wang, Haipeng Zeng, Yong Wang, Aoyu Wu 외

The modulation of voice properties, such as pitch, volume, and speed, is crucial for delivering a successful public speech. However, it is challenging to master different voice modulation skills. Though many guidelines a…

Sentence

Referee: Towards reference-free cross-speaker style transfer with low-quality data for expressive speech synthesis

2021-09-08 · Songxiang Liu, Shan Yang, Dan Su, Dong Yu

Cross-speaker style transfer (CSST) in text-to-speech (TTS) synthesis aims at transferring a speaking style to the synthesised speech in a target speaker's voice. Most previous CSST approaches rely on expensive high-qual…

Expressive Speech SynthesisSentenceSpeech SynthesisStyle Transfer+2

VoiceAssistant-Eval: Benchmarking AI Assistants across Listening, Speaking, and Viewing

2025-09-26 · Ke Wang, Houxing Ren, Zimu Lu, Mingjie Zhan 외 arxiv

The growing capabilities of large language models and multimodal systems have spurred interest in voice-first AI assistants, yet existing benchmarks are inadequate for evaluating the full range of these systems' capabili…

Rhythm Modeling for Voice Conversion

2023-07-12 · Benjamin van Niekerk, Marc-André Carbonneau, Herman Kamper

Voice conversion aims to transform source speech into a different target voice. However, typical voice conversion systems do not account for rhythm, which is an important factor in the perception of speaker identity. To …

RhythmVoice Conversion